> For the complete documentation index, see [llms.txt](https://cmo-ci.gitbook.io/access-quality-control-v1/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://cmo-ci.gitbook.io/access-quality-control-v1/coverage-vs-gc-content.md).

# Coverage vs GC content

Awareness of possible loss of accuracy in downstream sequencing results due to coverage bias

![](https://2763969089-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-M52gq1rRSDQOKMQGEuR%2F-M9Z39Nd7c6cJRknLQBm%2F-M9Z5pOs88Gy42ZQqJsH%2FScreen%20Shot%202020-06-11%20at%2011.59.00%20AM.png?alt=media\&token=eb3b0e3f-cafb-42ee-903a-3440f13ee71a)

**Theoretical Method**

Bin GC content of each region in the bam file into 5% intervals, and plot mean coverage across all regions that fall into each bin.&#x20;

**Technical Methods**

* Tool Used:
  * Waltz CountReads
  * aggregate\_bam\_metrics.sh
  * tables\_module.py
  * plots\_module.r
* Input
  * Standard bam
  * Collapsed unfiltered bam
  * ACCESS pool A bed file
* Output
  * sample\_id-intervals.txt

**Interpretations**\
Extreme base compositions, i.e., GC-poor or GC-rich sequences, lead to an uneven coverage or even no coverage of reads across the genome. This can affect downstream small variant and copy number calling. Both of which rely on consistent sequencing depth across all regions. Ideally this plot should be as flat as possible. The above example depicts a slight decrease in coverage at really high GC-rich regions, but is a good result for ACCESS.&#x20;
