As I see it now, a basic unit of information in CADB is a CA amino acid sequence (that belong to certain organism, family, subgroup, etc.) and for most CAs this info can be taken from Uniprot. To be specific, let's consider all CAs of mouse from UniProt and particularly CA with accn. Q64444. The question is what genomic information (chromosome num., exons, introns, transcripts, etc.) for Q64444 has to be in the CADB? Which source is better to use for getting genomic-related data? Ensembl?
Then, regarding what we discussed on meeting last time: coordinates. In case of Q64444 (and assuming we agree on taking coordinates from Ensembl) we have the following 'Ensembl' coordinates: Chromosome 11, 84,771,290-84,779,546. As I understood Csaba, these coordinates may be changed in next Ensembl releases. In fact, such changes are ok but this means that we need to have some routine that would be run once in a while to check/update coordinates (as well as, likely, something else). So, what kind of solution on coordinates may we have?
Thursday, October 2, 2008
Subscribe to:
Posts (Atom)