Sequence assembly from corrupted shotgun reads

Ganguly, S; Mossel, E; Racz, Miklos Z

Sequence assembly from corrupted shotgun reads

Author(s): Ganguly, S; Mossel, E; Racz, Miklos Z

Download

To refer to this page use: http://arks.princeton.edu/ark:/88435/pr1ps25

Full metadata record

DC Field	Value	Language
dc.contributor.author	Ganguly, S	-
dc.contributor.author	Mossel, E	-
dc.contributor.author	Racz, Miklos Z	-
dc.date.accessioned	2021-10-11T14:17:53Z	-
dc.date.available	2021-10-11T14:17:53Z	-
dc.date.issued	2016-08-10	en_US
dc.identifier.citation	Ganguly, S, Mossel, E, Racz, MZ. (2016). Sequence assembly from corrupted shotgun reads. IEEE International Symposium on Information Theory - Proceedings, 2016-August (265 - 269. doi:10.1109/ISIT.2016.7541302	en_US
dc.identifier.issn	2157-8095	-
dc.identifier.uri	http://arks.princeton.edu/ark:/88435/pr1ps25	-
dc.description.abstract	© 2016 IEEE. The prevalent technique for DNA sequencing consists of two main steps: shotgun sequencing, where many randomly located fragments, called reads, are extracted from the overall sequence, followed by an assembly algorithm that aims to reconstruct the original sequence. There are many different technologies that generate the reads: widely-used second-generation methods create short reads with low error rates, while emerging third-generation methods create long reads with high error rates. Both error rates and error profiles differ among methods, so reconstruction algorithms are often tailored to specific shotgun sequencing technologies. As these methods change over time, a fundamental question is whether there exist reconstruction algorithms which are robust, i.e., which perform well under a wide range of error distributions. Here we study this question of sequence assembly from corrupted reads. We make no assumption on the types of errors in the reads, but only assume a bound on their magnitude. More precisely, for each read we assume that instead of receiving the true read with no errors, we receive a corrupted read which has edit distance at most ϵ times the length of the read from the true read. We show that if the reads are long enough and there are sufficiently many of them, then approximate reconstruction is possible: we construct a simple algorithm such that for almost all original sequences the output of the algorithm is a sequence whose edit distance from the original one is at most O(ϵ) times the length of the original sequence.	en_US
dc.format.extent	265 - 269	en_US
dc.language.iso	en_US	en_US
dc.relation.ispartof	IEEE International Symposium on Information Theory - Proceedings	en_US
dc.rights	Author's manuscript	en_US
dc.title	Sequence assembly from corrupted shotgun reads	en_US
dc.type	Journal Article	en_US
dc.identifier.doi	doi:10.1109/ISIT.2016.7541302	-
pu.type.symplectic	http://www.symplectic.co.uk/publications/atom-terms/1.0/conference-proceeding	en_US

Files in This Item:

File	Description	Size	Format
Sequence assembly from corrupted shotgun reads.pdf		521.05 kB	Adobe PDF	View/Download

Show Simple Item Record