加载中…
个人资料
铁汉1990
铁汉1990
  • 博客等级:
  • 博客积分:0
  • 博客访问:1,333,641
  • 关注人气:828
  • 获赠金笔:0支
  • 赠出金笔:0支
  • 荣誉徽章:
相关博文
推荐博文
正文 字体大小:

【T】每日一生信--scaffold和contigs的区别

(2014-05-09 23:38:19)
标签:

scaffold

contigs

杂谈

分类: bioinformatic
该博文已整理到新地址:http://qinqianshan.com/ngs/bin/scaffold-contig/

A contig is a contiguous length of genomic sequence.
A scaffold is composed of contigs and gaps. Gap length can be guessed by incorporating information from paired ends or mate pairs

What is a Scaffold?
A scaffold is a portion of the genome sequence reconstructed from end-sequenced whole-genome shotgun clones. Scaffolds are composed of contigs and gaps. A contig is a contiguous length of genomic sequence in which the order of bases is known to a high confidence level. Gaps occur where reads from the two sequenced ends of at least one fragment overlap with other reads in two different contigs (as long as the arrangement is otherwise consistent with the contigs being adjacent). Since the lengths of the fragments are roughly known, the number of bases between contigs can be estimated.
【T】每日一生信--scaffold和contigs的区别

The goal of whole-genome shotgun assembly is to represent each genomic sequence in one scaffold; however, this is not always possible. One chromosome may be represented by many scaffolds (e.g., Chlamydomonas reinhardtii) or just a single scaffold (e.g., Human chromosome 19), depending on how completely the genome can be reconstructed, or assembled, from the available reads.  The relative locations of scaffolds in the genome are unknown.

Scaffolds are normally numbered approximately from largest to smallest. Some scaffolds may ultimately be filtered out of the assembly, resulting in skipped scaffold numbers.

In some cases, scaffolds can overlap. For example, in polymorphic genomes, regions with a high density of allelic differences between haplotypes may be split into separate sets of scaffolds, each representing one allele. Thus, a sequence that exists in only one location in the genome may appear on more than one scaffold.

Gaps are shown in the Genome Viewer as red lines or rectangles in the scaffold track (viewed in "full" mode). Contigs are shown in black. In FASTA sequences, gaps are represented by a series of Ns.

参考资料:
菜鸟的博客 http://blog.sina.com.cn/s/blog_4af3f0d20100kkkv.html
http://genome.jgi-psf.org/help/scaffolds.html

0

阅读 评论 收藏 转载 喜欢 打印举报/Report
  • 评论加载中,请稍候...
发评论

    发评论

    以上网友发言只代表其个人观点,不代表新浪网的观点或立场。

      

    新浪BLOG意见反馈留言板 欢迎批评指正

    新浪简介 | About Sina | 广告服务 | 联系我们 | 招聘信息 | 网站律师 | SINA English | 会员注册 | 产品答疑

    新浪公司 版权所有