Schema for Pfam in UCSC Gene - Pfam Domains in UCSC Genes
  Database: hg38    Primary Table: ucscGenePfam    Row Count: 62,551
Format description: Browser extensible data
fieldexampleSQL type info description
bin 585smallint(5) unsigned range Indexing field to speed chromosome range queries.
chrom chr1varchar(255) values Reference sequence chromosome or scaffold
chromStart 69168int(10) unsigned range Start position in chromosome
chromEnd 69972int(10) unsigned range End position in chromosome
name 7tm_4varchar(255) values Name of item
score 0int(10) unsigned range Score from 0-1000
strand +char(1) values + or -
thickStart 0int(10) unsigned range Start of where display should be thick (start codon)
thickEnd 0int(10) unsigned range End of where display should be thick (stop codon)
reserved 0int(10) unsigned range Used as itemRgb as of 2004-11-22
blockCount 1int(10) unsigned range Number of blocks
blockSizes 804,longblob   Comma separated list of block sizes
chromStarts 0,longblob   Start positions relative to chromStart

Sample Rows
 
binchromchromStartchromEndnamescorestrandthickStartthickEndreservedblockCountblockSizeschromStarts
585chr169168699727tm_40+0001804,0,
585chr169183698167TM_GPCR_Srsx0+0001633,0,
585chr169189699307tm_10+0001741,0,
586chr1187279195375WASH_WAHD0-000108,202,132,7,4,141,48,96,112,117,0,96,475,742,822,846,1159,1209,1511,7979,
588chr14507724515857tm_40-0001813,0,
588chr14509734515587tm_10-0001585,0,
590chr16857486865617tm_40-0001813,0,
590chr16859496865347tm_10-0001585,0,
592chr1943311943910SAM_10+000366,111,3,0,386,596,
592chr1943314943799SAM_20+000263,102,0,383,

Note: all start coordinates in our database are 0-based, not 1-based. See explanation here.

Pfam in UCSC Gene (ucscGenePfam) Track Description
 

Description

Most proteins are composed of one or more conserved functional regions called domains. This track shows the high-quality, manually-curated Pfam-A domains found in transcripts located in the UCSC Genes track.

Display Conventions and Configuration

This track follows the display conventions for gene tracks.

Methods

The sequences from the knownGenePep table (see UCSC Genes description page) are submitted to the set of Pfam-A HMMs which annotate regions within the predicted peptide that are recognizable as Pfam protein domains. These regions are then mapped to the transcripts themselves using the pslMap utility.

Credits

pslMap was written by Mark Diekhans at UCSC.

References

Finn RD, Mistry J, Tate J, Coggill P, Heger A, Pollington JE, Gavin OL, Gunasekaran P, Ceric G, Forslund K et al. The Pfam protein families database. Nucleic Acids Res. 2010 Jan;38(Database issue):D211-22. PMID: 19920124; PMC: PMC2808889