Thursday, February 19, 2009

Known Issues in STRING version 8.0

For each STRING version so far, only when we released it to the users did we find the last remaining bugs. Users often email us with their problems, and sometimes we are indeed to blame because there is an error. This is good (we think), because each bug found is a bug fixed - albeit only in the next release, usually.

So far, this is what we have found in release 8.0:

a) Some of our text-mining links do not show up in the corresponding evidence viewer. They are still correct, but the underlying text cannot be recovered and shown, for technical reasons. This happens because we developed a new feature that recognizes generic 'family' names for gene groups (like 'WNTs' for the various, homologous Wnt proteins). Within reasonable limits, such ambiguous names are now expanded to the individual protein members. However, we forgot to update the code of the text-viewer to reflect this ... we will do so in the next version.

b) Unfortunately, some of the prokaryotic genomes in this release are incomplete - in 43 cases we're missing a second (or third) minor chromosome. This was caused by a misunderstanding when parsing files from the RefSeq database: RefSeq provides an overview file that only lists one chromosome for each prokaryote, and we mistook that file for the full listing. Again, this will be fixed in the next release of STRING (on which we are already working). Obviously, we're now writing a new entry in our test suite that will prevent this type of error in the future - we will be checking the final gene counts of all organisms for consistency and also compare these counts against an external reference. Below is a list of affected organisms in the current release; if you're working with any of these, we recommend you continue using version 7.1 of STRING for now.

Luckily, no major model organisms are affected !!

Agrobacterium tumefaciens str. C58
Brucella abortus biovar 1 str. 9-941
Brucella melitensis 16M
Brucella melitensis biovar Abortus 2308
Brucella ovis ATCC 25840
Brucella suis 1330
Burkholderia ambifaria AMMD
Burkholderia cenocepacia AU 1054
Burkholderia cenocepacia HI2424
Burkholderia mallei ATCC 23344
Burkholderia mallei NCTC 10229
Burkholderia mallei NCTC 10247
Burkholderia mallei SAVP1
Burkholderia pseudomallei 1106a
Burkholderia pseudomallei 1710b
Burkholderia pseudomallei 668
Burkholderia pseudomallei K96243
Burkholderia sp. 383
Burkholderia thailandensis E264
Burkholderia vietnamiensis G4
Burkholderia xenovorans LB400
Deinococcus radiodurans R1
Haloarcula marismortui ATCC 43049
Leptospira borgpetersenii serovar Hardjo-bovis JB197
Leptospira borgpetersenii serovar Hardjo-bovis L550
Leptospira interrogans serovar Copenhageni str. Fiocruz L1-130
Leptospira interrogans serovar Lai str. 56601
Ochrobactrum anthropi ATCC 49188
Paracoccus denitrificans PD1222
Photobacterium profundum SS9
Pseudoalteromonas haloplanktis TAC125
Ralstonia eutropha H16
Ralstonia eutropha JMP134
Ralstonia metallidurans CH34
Rhodobacter sphaeroides 2.4.1
Rhodobacter sphaeroides ATCC 17029
Vibrio cholerae O1 biovar eltor str. N16961
Vibrio cholerae O395
Vibrio fischeri ES114
Vibrio harveyi ATCC BAA-1116
Vibrio parahaemolyticus RIMD 2210633
Vibrio vulnificus CMCP6
Vibrio vulnificus YJ016

That's it for known issues so far. But, do keep those emails coming - the feedback is very valuable !!

Thursday, January 29, 2009

Homology correction of co-occurrence and text-mining scores (updated)

[Update 29/01/2009: The homology correction is also applied to the text-mining channel starting with STRING 8.]

In order to avoid that gene duplications lead spurious functional associations, homologous proteins are down-weighed in the co-occurrence and text-mining channels. You will notice this on the score summary page of a link and if you have our SQL dumps.

Here's an example: The co-occurrence view looks fine for this pair of proteins.
However, the total score of 0.204 is less than the co-occurrence score:

The reason for this is that the proteins have some sequence similarity and are therefore down-weighted according to this formula:

effective co-occurrence score = co-occurrence score * (1 - homology score)

(The homology score is calculated from the bit score of the alignment.) In this case:

0.204 = 0.478 * ( 1 - 0.572 )


Thursday, January 15, 2009

Linking to individual networks (updated)

If you want to link to STRING or STITCH from your website, you can use the following URLs for simple queries:

For STITCH, you can use names of chemicals:

http://stitch.embl.de/interactions/aspirin?species=9606

You can also use identifiers, e.g. SwissProt or ATC codes:

http://string.embl.de/interactions/DRD1_HUMAN
http://stitch.embl.de/interactions/A01AD05?species=9606

(The 9606 specifies that you want human interactions, see NCBI taxonomy.)

Update: You can also link to networks with multiple items. As STRING saves the user's preference for proteins/COG mode, it's better to specify the target mode.

http://string.embl.de/interactionsList/zgc:73075%0Dzgc:136854?targetmode=proteins
http://string.embl.de/interactionsList/zgc:73075%0Dzgc:136854?targetmode=cogs
http://string.embl.de/interactionsList/KOG0044%0DKOG3656?targetmode=cogs

http://stitch.embl.de/interactionsList/DRD1_HUMAN%0Dpergolide?species=9606

You construct the URL by concatenating the protein names with "%0D" or "%0A" (an encoded carriage return / newline character).

Wednesday, January 14, 2009

New Year - New Major Release

Looks like 2009 will bring a lot of changes for both STRING and STITCH - and to lead the way, STRING has now been upgraded to version 8.0 !

This has been a major upgrade, and it has been some time in the making. We have almost doubled the number of organisms (again), and re-imported all the various pathways, protein-complexes and text-collections. We've also worked a lot behind the scenes, solidifying the API, further automating our data import and updating the way we display orthologous groups, to name just a few examples. All this has been possible only, really, because of our new sponsor - the Swiss Institute of Bioinformatics (SIB). Thanks guys !

More info about this new release is also available from here.

Monday, July 21, 2008

High-resolution images

We recently implemented a way to export high-res images (300 dpi, click the image below to see an example). This feature will go public with STRING 8 / STITCH 2, but if you're now using STRING or STITCH and want to prepare an image for publication, please get in touch with us (mkuhn embl de) and we can send you the image.

Thursday, June 26, 2008

How we compute scores (Part 1: experiment channel)

This is in response to a a question that we get quite frequently.
Sorry, it's a bit long - but this way it should contain sufficient detail to roughly understand how our scores come about (for the 'experiments channel' at least, and limited to protein mode). Have fun reading !

Christian von Mering (and Lars Jensen).


Procedure to compute experimental scores

  • first, we import information about which proteins have been shown to interact experimentally, from the following databases: INTACT, MINT, GRID, BIND, and DIP. To a small extent, this also includes experimental data that is not necessarily indicative of a direct physical interaction, such as genetic interaction data. Most, however, are from more-or-less direct, physical detection methods.
  • then, we map the proteins mentioned in these database onto the proteins in the STRING database - using identifiers, or (if needed) sequences.
  • next, we group all interactions by their supporting publication (PMID), and make them non-redundant (they might be reported under the same PMID from several databases). We also expand pulldowns of entire protein complexes using the 'spoke' model (i.e. assuming binary interactions from the tagged/immunoprecipitated protein to all of its co-purified partners).
  • then, we subdivide all interactions into 'small-scale, medium-scale, and high-throughput', based on the number of interactions reported by a single publication. These three classes are delineated by the extent of overlap with benchmark information, see below.
  • next, for each of these classes, we determine their 'reliability', by comparing them to our KEGG-benchmark. Briefly, an interaction between two proteins is counted as 'correct', when they are both annotated together in at least one 'KEGG-map', i.e. in at least one functional process / pathway. It is counted as 'incorrect', when the two proteins are annotated in KEGG, but never in the same pathway. Note that proteins that are not annotated at all in KEGG are not considered here).
  • for the small-scale experiments, which are only very few (per paper), we cannot benchmark each paper separately. Therefor, all such papers are lumped, benchmarked together, and we usually find them to be of quite high quality. As a result, we fix their score to some high number, for example 0.900 in the case of STRING version 7.1
  • for the medium-scale experiments, a separate score is computed for each publication, in a similar manner (some publications are found to report data of better reliability, other of lower reliability).
  • for the high-throughput experiments (there are less than 20 of these currently), we have enough information to be even a bit more specific: for each interaction in these sets, we can compute a 'raw score' from the data, because there are so many measurements done. Usually, this would be a score that describes how often a measurement has been confirmed, or how specific a particular interaction is, given the occurence of the two protein elsewhere throughout the dataset. These 'raw scores' are then binned, and each bin benchmarked separately, to arrive at a 'calibration curve', again using the KEGG pathways as a benchmark as described above. Thus, for these large sets, some interactions get a higher score, and others a lower score, depending on the information in the entire dataset.
  • this brings us to the cutoffs that determine whether something is small-scale, medium or high-throughput. This is defined on how many 'true-positives' are in the dataset: To be a large-scale dataset, we require at least 50 true positive interactions to enable the benchmarking. Otherwise, more than 20 true positive interactions will make it a medium-scale dataset, and the rest is small-scale. ("true positives" are defined as interactions where both proteins are in KEGG, and are sharing at least one KEGG map).
  • then, we have to deal with interactions supported by more than one independent dataset (i.e. by more than one publication). For those, the scores are 'added up'. Of course, they are not literally added up, but rather in a probabilistic integration, like so:
new_score = 1 - (1 - score_a) * (1 - score_b).

  • and finally, whe have to deal with interactions that are reported in multiple organisms, or in an organism other than the one of interest. This is called 'interaction transfer', and is a very important step to increase coverage. It is described in the 2005 STRING paper. Essentially, the better the orthology situation can be delineated (i.e. clear orthologs for both interacting partners can be identified), the bigger the score fraction that is transferred. Transferred interactions are integrated probabilistically as mentioned above, and interactions that are reported in two very similary organisms (say, mouse and rat), are considered redundant and transferred only once. Note that transferred scores are stored separately from the 'direct' scores in the database, so that all the transferred information can be discarded, if desired.

Tuesday, June 17, 2008

Downtime Wednesday morning

The STRING server will get a new disk tomorrow morning (European time), so there will be a downtime for STRING/STITCH. We hope everything will be working again in the early afternoon.

Update: We're back online, with enough room for the next version of STRING.