CYP2D6 copy number is the count of CYP2D6 gene copies present in a patient's genome — and the single most consequential lab decision when that count exceeds two is confirming which allele carries the duplication before assigning an activity score. A duplication of CYP2D6**1 creates an ultrarapid metabolizer; a duplication of CYP2D6**4 does not. Without allele-specific resolution, a total copy number of three is clinically uninterpretable. Every report your lab issues on a sample with copy-number variation (CNV) should reference three anchoring resources:
- PharmVar for current star-allele definitions and structural variant annotation conventions
- CPIC for activity score thresholds and phenotype-to-drug guidance
- AMP PGx Working Group recommendations for tiered clinical testing standards
These are not optional citations. They are the evidentiary backbone of a defensible report.
Key Takeaways
CYP2D6 copy number is clinically meaningful only when the functional status of the duplicated allele is known — total CN alone cannot support a defensible phenotype assignment.
| Point | Details |
|---|---|
| Allele identity determines phenotype | A duplication of *1 or *2 creates an ultrarapid metabolizer; a duplication of *4 does not change phenotype from poor metabolizer. |
| 12.6% of US samples carry CNV | A 31,563-sample cohort found ~12.6% with zero, one, or three or more copies, justifying routine CNV testing in clinical panels. |
| Use at least two target regions | Exon 9-only assays produce false-positive CN gain in CYP2D7 hybrid samples; add a 5' or intron 6 target to every CN assay design. |
| Reflex ambiguous calls | Any CN result outside the normal two-copy range, or any discordance between target regions, requires orthogonal confirmation before phenotype assignment. |
| Cite PharmVar, CPIC, and AMP PGx | Every clinical report should reference the allele definition source, the activity score framework, and the guideline version used for phenotype mapping. |
Table of Contents
- How CYP2D6 copy-number variation changes drug response
- Why the CYP2D6 locus is so difficult to genotype accurately
- Comparing laboratory methods for CYP2D6 copy-number determination
- Validation, quality controls, and reference materials for CYP2D6 CNV testing
- Converting copy-number calls into star-alleles, activity scores, and report language
- Practical troubleshooting when CN calls disagree or are ambiguous
- Long-read sequencing and emerging callers: what's ready for your lab now
- A lab director's perspective on CYP2D6 CNV in routine PGx panels
- Sources
- FAQ
How CYP2D6 copy-number variation changes drug response
CNV can shift a patient's metabolic phenotype across the full spectrum — from poor metabolizer (PM) to ultrarapid metabolizer (UM) — depending entirely on the functional status of the allele that is duplicated or deleted. A deletion of both copies (*CYP2D6**5/*5) eliminates enzyme activity and produces a PM. A duplication of a fully functional allele (*1xN or *2xN) increases activity and can produce a UM. The clinical consequences of each extreme are distinct and, in some cases, life-threatening.
Drugs with the highest clinical stakes for CYP2D6 CNV include:
- Opioids (codeine, tramadol, hydrocodone): Codeine is a prodrug requiring CYP2D6-mediated conversion to morphine. UMs convert codeine to morphine at an accelerated rate, creating toxicity risk — including respiratory depression and death in pediatric patients. PMs receive no analgesic benefit. The FDA biomarker labeling for codeine explicitly references CYP2D6 phenotype. For a detailed prescribing walkthrough, see codeine and CYP2D6 prescribing guidance.
- Tamoxifen: CYP2D6 converts tamoxifen to endoxifen, its active metabolite. PMs and intermediate metabolizers (IMs) produce substantially less endoxifen, which may reduce efficacy in hormone receptor-positive breast cancer.
- Tricyclic antidepressants (amitriptyline, nortriptyline, imipramine): PMs accumulate parent drug and active metabolites, increasing cardiotoxicity and anticholinergic adverse effects. CPIC provides dosing guidance for this class.
- Atomoxetine: PMs experience four- to tenfold higher plasma exposure compared to normal metabolizers, requiring dose reduction.
- Select SSRIs (fluoxetine, paroxetine, fluvoxamine): These are both substrates and inhibitors of CYP2D6, making CNV-driven phenotype shifts clinically relevant for both efficacy and drug-drug interaction risk.
- Antiemetics (ondansetron, metoclopramide): CYP2D6 activity affects exposure and, for metoclopramide, extrapyramidal risk.
A 31,563-sample US clinical cohort found that approximately 12.6% of samples carried CNV — zero, one, or three or more copies — with roughly 5.2% showing three or more copies. Not all of those high-copy samples represent functional UMs, because the allele content of the duplication was not always determined. That gap between detected CN and confirmed phenotype is precisely where laboratory rigor matters most.
CPIC translates genotype into phenotype through an activity score (AS) system: each allele is assigned a value (1.0 for fully functional, 0.5 for reduced function, 0.0 for nonfunctional), the two allele scores are summed, and the total maps to a phenotype category. A duplication multiplies the score of the duplicated allele accordingly — *2xN with two copies contributes 2.0 to the AS, while *4xN contributes 0.0 regardless of copy count. Reporting a total CN without an allele-specific call leaves the activity score calculation incomplete.
Why the CYP2D6 locus is so difficult to genotype accurately
Short-read NGS reads from CYP2D7 routinely map to CYP2D6 reference coordinates, inflating apparent copy number or masking real variants. PCR primers designed against CYP2D6 can co-amplify CYP2D7 unless carefully validated. The Frontiers pharmacology methods review identifies pseudogene homology as the primary reason clinical assays frequently fail to specify which allele is duplicated.
Common structural features that complicate CN interpretation:
- Whole-gene duplications and multiplications: The entire CYP2D6 locus is tandemly duplicated, sometimes producing three, four, or more copies. Annotation follows the xN convention (*1xN, *2xN), but as PharmVar's gene-support documentation notes, duplicated copies are typically assumed identical even though sequencing both copies is rarely performed — an assumption that may be incorrect.
- *Whole-gene deletion (5): A recombination event deletes the entire gene, yielding zero copies on that chromosome.
- **CYP2D6-CYP2D7 hybrid alleles (*36, *68, 13, 61, etc.): Recombination between CYP2D6 and CYP2D7 produces chimeric genes. *36 contains CYP2D7 sequence in exon 9, which is why exon 9-only copy-number assays generate false-positive CN gain results in samples carrying this hybrid. The US CNV distribution study documented this failure mode explicitly.
- Gene conversion events: Segments of CYP2D7 sequence are introduced into an otherwise intact CYP2D6 gene, altering function without changing total copy count.
The practical consequence for your lab: when a sample returns more than two copies, treat the result as provisional until you have confirmed the structural basis. Exon 9-only assays are insufficient for that confirmation. Phasing — determining which chromosome carries the extra copy — requires either allele-quantifying methods or long-range structural analysis.
Comparing laboratory methods for CYP2D6 copy-number determination
Choose the method that matches your lab's requirement for allele-specific resolution against its throughput, cost, and validation maturity constraints. No single method solves every problem; most high-volume clinical labs use a tiered approach.

| Method | Total CN vs. Allele-Specific | Hybrid/Pseudogene Susceptibility | Throughput / Cost | DNA Input | Turnaround | Clinical Validation Maturity |
|---|---|---|---|---|---|---|
| qPCR / TaqMan CNV assay | Total CN only | High (exon 9 false positives) | High / Low | 10–50 ng | 1–2 days | Well-established; CLIA-validated in many labs |
| MLPA (MRC Holland) | Total CN; some allele inference | Moderate (probe-specific) | Medium / Medium | 50 ng | 2–3 days | Validated; used in CAP PT programs |
| ddPCR / dPCR (Bio-Rad QX200, Absolute Q) | Total CN; allele-specific with allele-discriminating probes | Moderate | Medium / Medium-High | 10–50 ng | 1–2 days | Growing clinical evidence; CLIA-implementable |
| Pyrosequencing allele quantification (Qiagen PSQ) | Allele-specific ratio | Low (sequence-based discrimination) | Low-Medium / Medium | 50–100 ng | 2–3 days | Validated vs. Coriell and TaqMan controls |
| Short-read NGS + specialized callers (Cyrius, Stargazer) | CN field + phasing inference | High (read misalignment) | High / Low-Medium | 100 ng | 3–7 days | Validated in research cohorts; flags uncertain calls |
| XL-PCR + Sanger or short-read sequencing | Allele-specific (structural) | Low (amplicon-specific) | Low / High | 100 ng | 3–5 days | Gold standard for structural confirmation |
| Long-read targeted sequencing (nCATS / nanopore) | Fully allele-specific + phasing | Very Low | Low / High | 1–5 µg | 5–10 days | Research-stage; clinical validation in progress |

Target regions interrogated by method:
Method-specific notes:
qPCR / TaqMan (Thermo Fisher TaqMan Copy Number Assays) are the workhorse for high-volume labs. They are fast, inexpensive, and well-understood, but they report total CN only and are vulnerable to exon 9 false positives from CYP2D7 hybrids. Running at least two target regions — one 5' (upstream or intron 2) and one 3' (intron 6 or exon 9) — is the minimum acceptable design. When the two regions disagree, reflex to an orthogonal method.
MLPA (MRC Holland) interrogates multiple probes across the gene in a single reaction, providing better structural resolution than single-target qPCR. It can detect partial gene duplications and deletions and is compatible with CAP proficiency testing programs. Per-sample cost is moderate, and the method is well-suited to low-to-medium throughput labs.
ddPCR / dPCR (Bio-Rad QX200 Droplet Digital PCR, Absolute Q) offers absolute quantification without a standard curve, improving precision for borderline CN calls (e.g., distinguishing two from three copies). With allele-discriminating probe pairs, ddPCR can provide allele-specific ratios in heterozygous samples. The Frontiers dPCR methods insight recommends prioritizing dPCR or long-read sequencing when phase-specific characterization is clinically required.
Pyrosequencing allele quantification (Qiagen PSQ platform) uses sequencing-by-synthesis to measure allele ratios at informative SNP positions, enabling allele-specific CN inference in heterozygous samples. A validated method demonstrated 100% concordance with TaqMan and Coriell reference controls in distinguishing two, three, and four copy states. Throughput is lower than qPCR, making it better suited as a reflex or confirmation method than a primary screen.
Short-read NGS with Cyrius or Stargazer can call CN from read-depth data and include copy_number fields and filter flags for low-confidence calls. Both tools handle many common configurations accurately but flag high-CN or hybrid configurations for manual review. The All of Us pharmacogenomics pipeline uses Cyrius for star-allele calling at scale, which demonstrates research-cohort feasibility, though clinical labs must independently validate these callers under CLIA conditions.
XL-PCR + sequencing remains the structural confirmation gold standard. Long-range amplicons spanning the CYP2D6/CYP2D7 junction can distinguish hybrid alleles from true duplications and provide sequence-level evidence for allele identity. Throughput is low and cost is high, so this method is best reserved for samples where other methods return discordant or ambiguous results.
Validation, quality controls, and reference materials for CYP2D6 CNV testing
Validation must demonstrate accuracy against characterized reference materials and orthogonal methods for both total CN and allele-specific calls. A CN assay that performs well on two-copy samples but has not been challenged with characterized three-copy, deletion, and hybrid samples does not meet CLIA or CAP expectations for a clinical test.
Required reference materials and controls:
- Coriell Cell Repositories: Characterized cell lines with known CYP2D6 genotypes, including deletion homozygotes, duplications, and hybrid alleles. These are the standard positive controls for CN validation in US clinical labs.
- GeT-RM (Genetic Testing Reference Materials Coordination Program): CDC-coordinated characterized samples with consensus genotype calls from multiple laboratories. GeT-RM materials provide independent confirmation of allele assignments and are particularly valuable for validating rare or complex configurations.
- Positive CN controls: At minimum, include samples representing CN = 0 (**5/*5), CN = 1 (deletion heterozygote), CN = 2 (normal), and CN ≥ 3 (duplication) in every validation run.
- Negative controls: No-template controls and samples with known hybrid alleles that should not trigger a duplication call in a correctly designed assay.
Validation checklist for CLIA/CLSI compliance:
- Analytical sensitivity and specificity across the full CN range (0–4+ copies)
- Reproducibility: intra-run, inter-run, and across DNA extraction methods and input concentrations
- Limit of detection for multiplications (can the assay reliably distinguish CN = 3 from CN = 4?)
- Assay linearity across the expected CN range
- Interference testing: performance in samples with known CYP2D7/CYP2D8 hybrid alleles
- Orthogonal confirmation plan: define in your SOP which results trigger reflex to a second method
QC best practices in routine operation:
- Run at least two target regions per sample (5' and 3' targets) to detect discordant signals
- Include batch controls with every run; track CN ratios over time to detect reagent drift
- Review all software filter flags from Cyrius or Stargazer before releasing results — do not auto-release flagged calls
- Log uncertain calls explicitly in the LIS with a notation indicating the basis for uncertainty and the reflex plan
The AMP PGx Working Group and CPIC recommendations provide tiered guidance for allele inclusion and technical feasibility, and they stress the use of validated analytical approaches and reference sequences. Your validation documentation should cite the specific reference assembly used — NG_008376.4 (LRG_303) or GRCh38 coordinates — because star-allele definitions in PharmVar are tied to specific reference sequences, and a mismatch between your annotation and the current PharmVar reference is a reportable discrepancy.
Pro Tip: Run GeT-RM samples blind alongside your validation cohort rather than as labeled positive controls. Blind testing reveals assay failure modes that labeled controls can mask, particularly for hybrid alleles where the expected result is already known to the analyst.
Converting copy-number calls into star-alleles, activity scores, and report language
Report diplotypes with allele-specific CN where known, show the activity score calculation explicitly, and provide phenotype plus actionable guidance or flags. A total CN of three with no allele assignment is not a complete clinical result.
Example report elements:
- Measured total CN (e.g., CN = 3 by TaqMan, confirmed by MLPA)
- Called diplotype with CN notation: **1/*2xN (where N = 2) or, when phasing is unresolved, (**1/**2) x2 with a notation that allele-specific assignment is inferred
- Activity score calculation: *1 (AS = 1.0) + *2xN (AS = 1.0 × 2 = 2.0) = total AS 3.0
- Phenotype mapping per CPIC: AS ≥ 2.25 = Ultrarapid Metabolizer
- Clinical interpretation: drug-specific guidance or flag referencing CPIC guideline version and date
- Orthogonal confirmation status: confirmed by [method] or "allele-specific CN not determined; orthogonal testing recommended"
Diplotype to activity score and phenotype mapping:
| Diplotype Example | Activity Score | CPIC Phenotype |
|---|---|---|
| **1/*1 | 2.0 | Normal Metabolizer |
| **1/*4 | 1.0 | Intermediate Metabolizer |
| **4/*4 | 0.0 | Poor Metabolizer |
| **5/*5 | 0.0 | Poor Metabolizer |
| **1/*2xN (N=2) | 3.0 | Ultrarapid Metabolizer |
| **2xN/*4 (N=2) | 2.0 | Normal Metabolizer |
| **36/*10 | 0.0 | Poor Metabolizer |
When only total CN is known and allele-specific assignment is not possible, use language such as: "Total CYP2D6 copy number = 3 by [method]. Allele-specific copy number not determined. Activity score and phenotype assignment are provisional pending orthogonal confirmation. Consider reflex testing to resolve allele-specific duplication." This phrasing satisfies the PharmVar reporting guidance requirement to reflect ambiguity rather than assign a false-confidence phenotype.
For genotype-to-phenotype mapping workflows in your reporting pipeline, the activity score must be calculated from allele-specific data, not from total CN alone.
Pro Tip: *When reporting a duplicated allele, always state the evidence basis for the allele identity call — e.g., "duplication allele identified as 2 by pyrosequencing allele quantification concordant with Coriell NA17102." This creates an auditable chain from raw CN to phenotype that satisfies both CLIA documentation requirements and clinically defensible PGx report standards.
Practical troubleshooting when CN calls disagree or are ambiguous
Ambiguous CN calls require stepwise orthogonal testing focused on allele-specific evidence. Do not assign a definitive phenotype from a single ambiguous assay result.
Stepwise resolution protocol:
- Step 1 — Re-extract and re-run: Confirm the discordant result is not a sample quality or pipetting artifact. Check DNA concentration, A260/A280, and fragmentation before proceeding.
- Step 2 — Test alternative target regions: If the initial call used exon 9 only, add a 5' upstream or intron 2 target. Discordance between 5' and 3' CN signals is a strong indicator of a hybrid allele.
- Step 3 — Apply allele-quantifying methods: Pyrosequencing at an informative SNP position or ddPCR with allele-discriminating probes can resolve whether the extra copy is on the same chromosome as a known variant allele.
- Step 4 — XL-PCR + sequencing for structural resolution: When steps 1–3 remain inconclusive, long-range PCR spanning the CYP2D6/CYP2D7 junction provides sequence-level structural evidence.
- Step 5 — Escalate to long-read phasing or reference lab: If your lab cannot achieve allele-specific resolution internally, send the sample to a reference laboratory with validated long-read or XL-PCR capabilities.
Common pitfalls and fixes:
- **Exon 9 false positives from CYP2D7 hybrids (36, 68): Add intron 6 or 5' upstream target to the assay. Discordance between exon 9 (CN = 3) and intron 6 (CN = 2) signals hybrid presence.
- ***36/10 confusion: *36 contains CYP2D7 exon 9 sequence and is often found in tandem with *10. This configuration can appear as a duplication on exon 9-based assays while the patient is actually a PM. Confirm with SNP genotyping at the *10 position (c.100C>T) and structural analysis.
- Sample fragmentation causing read-depth bias in NGS: Degraded DNA produces uneven read depth that mimics CN gain or loss. Check fragment size distribution before NGS library preparation; set a minimum input quality threshold in your SOP.
- Batch or pipetting variance in qPCR: CN ratios near integer boundaries (e.g., 2.4–2.6) are unreliable. Set a reflex threshold — for example, any CN ratio outside 0.7–1.3 per copy unit triggers orthogonal confirmation.
Long-read sequencing and emerging callers: what's ready for your lab now
Targeted long-read sequencing and improved bioinformatic phasing are the most promising paths to routine allele-specific CN resolution, though most approaches remain in the clinical validation stage for US labs.
nCATS and Cas9-targeted nanopore sequencing use CRISPR-Cas9 to enrich for the CYP2D6 locus before Oxford Nanopore sequencing, enabling long reads that span the entire gene and its flanking regions. The nCATS + CoLoRGen pipeline has demonstrated the ability to reconstruct CYP2D6-CYP2D7 hybrid alleles and unambiguously assign CN to specific haplotypes when sufficient on-target depth is achieved. Current barriers include variable on-target yield and higher per-sample cost compared to qPCR or MLPA. For labs considering adoption, a parallel testing phase against characterized Coriell and GeT-RM samples is the minimum validation entry point.
Cyrius and Stargazer are the two most widely used short-read callers for CYP2D6 star-allele assignment from NGS data. Both produce a copy_number field alongside the star-allele call and include filter flags for configurations where the call confidence is low. Cyrius is used in the All of Us Research Program's pharmacogenomics pipeline, which provides a large-scale research validation context. Neither tool eliminates the need for manual review of flagged calls, and neither replaces orthogonal confirmation for high-CN or hybrid configurations in a clinical setting.
Cyrius and Stargazer are appropriate primary callers for standard diplotypes in high-throughput NGS workflows, but any call carrying a low-confidence flag or involving CN ≥ 3 should be treated as provisional until allele-specific evidence is obtained from an orthogonal method. Automated pipelines that auto-release flagged high-CN calls without review are a known source of phenotype misassignment.
Adoption guidance for US clinical labs:
- Pilot phase: Run long-read or advanced caller methods in parallel with your validated primary assay on 50–100 characterized samples. Document concordance rates and failure modes.
- Parallel testing: For a defined period, report results from both methods and resolve discordances with XL-PCR or Coriell reference materials.
- Phased roll-out: Introduce the new method as the primary assay for specific indication types (e.g., samples with initial CN ≥ 3 by qPCR) before full panel deployment.
Pro Tip: Before committing to a long-read platform for CYP2D6, request on-target yield data from the vendor for the specific Cas9 guide RNA set you plan to use. Published nCATS papers report variable enrichment efficiency, and your lab's DNA quality profile (FFPE vs. blood-derived) will significantly affect yield.
A lab director's perspective on CYP2D6 CNV in routine PGx panels
The minimum acceptable testing strategy for a US clinical lab in 2026 is a multi-target CN assay covering at least two genomic regions (one 5' and one 3' target), orthogonal confirmation for any sample returning more than two copies, and PharmVar/CPIC-aligned reporting that explicitly states the allele-specific basis for every activity score calculation.
The trade-offs are real and worth naming directly:
- Cost vs. precision: A two-target qPCR assay costs a fraction of ddPCR or long-read sequencing per sample, but it cannot phase duplications. For labs running high volumes of opioid-relevant panels, the clinical and liability costs of a misassigned UM phenotype in a codeine patient almost always justifies the investment in a reflex confirmation workflow.
- Staffing and skill requirements: Pyrosequencing allele quantification and XL-PCR require molecular staff comfortable with method-specific troubleshooting. If your lab does not have that expertise, a validated ddPCR workflow with allele-discriminating probes is a more practical path to allele-specific resolution than pyrosequencing.
- Integration with PGx reporting pipelines: CN calls do not exist in isolation. They feed into diplotype assignments, activity scores, and ultimately the clinical report. Labs that have not built a structured PGx reporting pipeline often find that CN ambiguity propagates silently into phenotype calls because there is no systematic trigger for reflex testing.
My view is that long-read phasing should be prioritized for any patient population where the clinical stakes of a misassigned UM or PM phenotype are highest — specifically, pediatric patients receiving codeine or tramadol, patients initiating tamoxifen for breast cancer, and patients on high-dose tricyclics where PM-driven toxicity is a documented risk. For the broader population in a general PGx panel, a well-validated two-target qPCR with a defined reflex protocol is a defensible and practical standard.
Sources
- PharmVar Tutorial on CYP2D6 Structural Variation Testing and Recommendations on Reporting
- A rapid allele quantification-based pyrosequencing method for CYP2D6 copy number variation
- CYP2D6_Variation_v1.6
This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.
FAQ
What happens when the CYP2D6 gene is duplicated?
When a functional CYP2D6 allele is duplicated, the patient produces more CYP2D6 enzyme, which can shift their phenotype to ultrarapid metabolizer. This increases conversion of prodrugs like codeine to active metabolites at an accelerated rate, raising toxicity risk, while diminishing efficacy of drugs that require normal CYP2D6 activity for therapeutic effect.
How do labs determine CYP2D6 copy number?
The most common primary method is qPCR using TaqMan Copy Number Assays targeting at least two genomic regions. For allele-specific resolution or ambiguous results, labs reflex to ddPCR, pyrosequencing allele quantification, MLPA, or XL-PCR plus sequencing. Short-read NGS with callers like Cyrius or Stargazer can also call CN from read-depth data, though flagged results require orthogonal confirmation.
How common is it to be a poor metabolizer of CYP2D6?
Is Adderall metabolized by CYP2D6?
Amphetamine (the active component of Adderall) is partially metabolized by CYP2D6, though CYP2D6 is not the primary elimination pathway. CYP2D6 polymorphism impacts amphetamine exposure to a modest degree, and the FDA labeling notes the interaction, but CYP2D6 phenotype is not currently a primary clinical decision point for Adderall dosing in the way it is for codeine or atomoxetine.
When should a lab escalate to long-read sequencing for CYP2D6?
Escalate to long-read sequencing when standard methods return discordant CN signals across target regions, when a hybrid allele is suspected but cannot be confirmed by XL-PCR, or when the clinical stakes of a misassigned phenotype are high (e.g., pediatric opioid dosing, tamoxifen initiation). Long-read platforms like nCATS with nanopore sequencing provide allele-specific phasing that no short-read or qPCR method can match.
