Skip to main content

Identifying structural variants using linked-read sequencing data.

Author(s): Elyanow, Rebecca; Wu, Hsin-Ta; Raphael, Benjamin J

Download
To refer to this page use: http://arks.princeton.edu/ark:/88435/pr1sz6r
Abstract: MOTIVATION:Structural variation, including large deletions, duplications, inversions, translocations and other rearrangements, is common in human and cancer genomes. A number of methods have been developed to identify structural variants from Illumina short-read sequencing data. However, reliable identification of structural variants remains challenging because many variants have breakpoints in repetitive regions of the genome and thus are difficult to identify with short reads. The recently developed linked-read sequencing technology from 10X Genomics combines a novel barcoding strategy with Illumina sequencing. This technology labels all reads that originate from a small number (∼5 to 10) DNA molecules ∼50 Kbp in length with the same molecular barcode. These barcoded reads contain long-range sequence information that is advantageous for identification of structural variants. RESULTS:We present Novel Adjacency Identification with Barcoded Reads (NAIBR), an algorithm to identify structural variants in linked-read sequencing data. NAIBR predicts novel adjacencies in an individual genome resulting from structural variants using a probabilistic model that combines multiple signals in barcoded reads. We show that NAIBR outperforms several existing methods for structural variant identification-including two recent methods that also analyze linked-reads-on simulated sequencing data and 10X whole-genome sequencing data from the NA12878 human genome and the HCC1954 breast cancer cell line. Several of the novel somatic structural variants identified in HCC1954 overlap known cancer genes. AVAILABILITY AND IMPLEMENTATION:Software is available at compbio.cs.brown.edu/software. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Publication Date: Jan-2018
Citation: Elyanow, Rebecca, Wu, Hsin-Ta, Raphael, Benjamin J. (2018). Identifying structural variants using linked-read sequencing data.. Bioinformatics (Oxford, England), 34 (2), 353 - 360. doi:10.1093/bioinformatics/btx712
DOI: doi:10.1093/bioinformatics/btx712
ISSN: 1367-4803
EISSN: 1367-4811
Pages: 353 - 360
Language: eng
Type of Material: Journal Article
Journal/Proceeding Title: Bioinformatics (Oxford, England)
Version: Author's manuscript



Items in OAR@Princeton are protected by copyright, with all rights reserved, unless otherwise indicated.