> For the complete documentation index, see [llms.txt](https://cmo-ci.gitbook.io/access-quality-control-v1/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://cmo-ci.gitbook.io/access-quality-control-v1/umi-family-types-composition-pool-a.md).

# UMI Family types Composition (Pool A)

Understanding the relative abundance of each fragment subtype

![](https://2763969089-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-M52gq1rRSDQOKMQGEuR%2F-M9Z39Nd7c6cJRknLQBm%2F-M9Z6KW76yFf6MRZ-OgP%2FScreen%20Shot%202020-06-11%20at%2012.01.13%20PM.png?alt=media\&token=e0e006bd-054f-48fb-9cbd-01748573a2a6)

**Theoretical Method**

Marianas performs read grouping based on the 6-base UMI sequence (three from each side of the DNA fragment), as well as the fragment start position ~~(and stop position?)~~. If multiple read pairs have the same information for these two metrics, they will be grouped into the same UMI "family".&#x20;

UMI family types are defined by the following categories:

* Duplex: both top and bottom strand were found for this fragment
* Simplex: only one of (top|bottom) strand was sequenced, and >=3 copies for that strand were found
* Sub-Simplex: exactly 2 copies of a single strand were found
* Singletons: exactly 1 copy of a single strand was found

**Technical Methods**

* Tool Used:
  * Marianas
  * make\_umi\_qc\_tables.sh
  * plots\_module.r
* Input
  * Marianas collapsed fastqs
* Output
  * family-types-A.txt

**Interpretations**

Duplex families are valuable for their low noise rate after collapsing, thus we'd like to see as high of a duplex "saturation" as possible. If this value is lower, we may not have captured enough of the original molecules to find both strands after PCR replication.&#x20;
