Average region length	1.18
Description	This dataset contains data on DNA sequencing.<br>
	It is directly derived from data available on the old TCGA portal and transformed into GDM format through the TCGA2BED pipeline.<br>
	Documentation is available at http://bioinf.iasi.cnr.it/tcga2bed/data/TCGA2BED_format_definition.pdf.<br>
	<br>
	This type of next generation sequencing (NGS) experiment discovers mutations by aligning DNA sequences derived from tumor samples to sequences derived from normal samples and a reference sequence.<br>
	<br>
	The dataset includes tab separated BED files, in which the DNA-seq .maf file is converted, with the following fields:<br>
	<ol>
	<li>chrom (i.e., the name of the chromosome, e.g., chr3, chrY, chr2_random, retrieved from the 5. field of the TCGA maf file)</li>
	<li>chromStart (i.e., the starting position of the feature in the chromosome or scaffold, e.g., 999, retrieved from the 6. field of the TCGA maf file)</li>
	<li>chromEnd (i.e., the ending position of the feature in the chromosome or scaffold, e.g., 1000, retrieved from the 7. field of the TCGA maf file)</li>
	<li>strand (i.e., it defines the strand, either '+' or '-'., retrieved from the 8. field of the TCGA maf file)</li>
	<li>hugo_symbol (i.e., the symbol of the gene related to the reported variant, if it exists, e.g., "EGFR", retrieved from the 1. field of the TCGA maf file)</li>
	<li>entrez_gene_id (i.e., the Entrez gene ID of the gene related to the reported variant, if it exists, e.g., "1956", retrieved from the 2. field of the TCGA maf file)</li>
	<li>variant_classification (i.e., the classification of the reported variant, e.g., "Missense_Mutation", retrieved from the 9. field of the TCGA maf file)</li>
	<li>variant_type (i.e., the type of mutation, e.g., "INS", retrieved from the 10. field of the TCGA maf file)</li>
	<li>reference_allele (i.e., the plus strand reference allele at this position, e.g., "A", retrieved from the 11. field of the TCGA maf file)</li>
	<li>tumor_seq_allele1 (i.e., the tumor sequencing (discovery) allele 1, e.g., "C", retrieved from the 12. field of the TCGA maf file)</li>
	<li>tumor_seq_allele2 (i.e., the tumor sequencing (discovery) allele 2, e.g., "G", retrieved from the 13. field of the TCGA maf file)</li>
	<li>dbsnp_rs (i.e., the latest dbSNP rs ID, e.g., "rs12345", retrieved from the 14. field of the TCGA maf file)</li>
	<li>tumor_sample_barcode (i.e., the BCR aliquot barcode for the tumor sample, e.g., "TCGA-02-0021-01A-01D-0002-04", retrieved from the 16. field of the TCGA maf file)</li>
	<li>matched_norm_sample_barcode (i.e., the BCR aliquot barcode for the matched normal sample, e.g., "TCGA-02-0021-10A-01D-0002-04", retrieved from the 17. field of the TCGA maf file)</li>
	<li>match_norm_seq_allele1 (i.e., the Matched normal sequencing allele 1, e.g., "T", retrieved from the 18. field of the TCGA maf file)</li>
	<li>match_norm_seq_allele2 (i.e., the Matched normal sequencing allele 2, e.g., "ACGT", retrieved from the 19. field of the TCGA maf file)</li>
	<li>matched_norm_sample_uuid (i.e., the BCR aliquot UUID for matched normal, e.g., "567e8487-e29b-32d4-a716-446655443246", retrieved from the 34. field of the TCGA maf file)</li>
	</ol>
Download date	2017/10/04 0:0:0
Number of regions	1319236
Number of samples	6914
Size	285.94 MB
Upload date	2017/10/05 0:0:0