IGS Data Representation

From GMOD
Revision as of 06:16, 15 January 2009 by Jorvis (Talk | contribs)

Jump to: navigation, search

Chado is an elegant schema that can hold nearly anything from gene annotations to an MP3 collection. This fabulous flexibility comes with a price - different MODs arrive at different ways of storing the same biological information. This page is not meant to be a tutorial of how YOU should model your biological information in Chado. Rather, it is simply a brain dump of the way we are doing things at IGS, for better or worse.

The reference document is currently the Chado Best Practices page, into which much of this information may become merged at some point.

What we store

We currently use the Chado schema primarily to store genome annotation data, including comparative genomics. This includes both read-only databases from 'finished' annotations and ongoing, actively modified data. Prokaryotes and Eukaryotes are both represented in our datasets and we use Sybase, MySQL and PostgreSQL back-ends.

Structural annotation

Gene models

Canonical gene model

Functional annotation