AI assisted with source organization and draft preparation; a human editor verified the source links, claims, limitations, and final public wording.
What the paper claims
The authors report that reference pretraining followed by personalized DNA–RNA fine-tuning improved mean across-individual Spearman correlation for familiar genes, with rp-Caduceus reaching 0.194 in the cited 50-train/100-test comparison. They report near-zero DNA-only transfer to 287 genes held out on chromosomes 5 and 10. Adding matched ATAC-seq increased correlation on unseen genes, but an ATAC-only control performed similarly or slightly better. These are author-reported benchmark results, not an independent reproduction or clinical validation.
What was tested
The study used paired 1000 Genomes genotypes and GEUVADIS lymphoblastoid-cell RNA-seq for 3,259 genes selected for significant European cis-eQTLs. Reference training used 2,972 genes; personalized training used 200 genes outside chromosomes 5 and 10; 287 chromosome-5/10 genes were reserved as unseen loci. The main split was 50 training and 100 held-out individuals. Personal haplotypes substituted SNVs, but not indels, into hs37d5. Frozen embeddings from several genomic models were passed to a convolutional decoder and evaluated primarily by gene-wise Spearman correlation across held-out people. The study also used a 50/44 matched genotype–RNA–ATAC split, ancestry transfer tests, linear PrediXcan-style baselines, and ATAC-only controls.
What the evidence supports
The strongest supported conclusion is narrow: paired personal-genome and RNA training improved prediction across new people for genes represented during personalized training, although absolute correlations remained modest. The evidence did not establish reliable DNA-only generalization to unseen loci; increasing the number of training individuals did not remove that bottleneck. Observed ATAC-seq improved unseen-gene correlations, but the ATAC-only control indicates that measured chromatin accessibility itself explained much of the gain rather than demonstrating superior genotype-to-expression transfer.
Limitations
The primary personalized cohort was small, the data came from one immortalized lymphoblastoid-cell context, and there was no independent cohort, tissue, or prospective replication. The gene set was enriched for known European cis-eQTLs, the unseen-locus result used one fixed chromosome split, and transfer weakened in YRI individuals. Personal sequences omitted indels and structural variants. Main models used final-epoch selection without a general validation set. The code repository lacks a tagged release, a fully pinned environment, and some final-paper analysis scripts. The results are not suitable for clinical or reliable per-gene use.
Some detailed values are available only in figures or downloadable source data, and no independent reproduction was located. The ancestry associations are not causal explanations. “Few-shot” is the authors’ framing for 50 paired individuals across 200 genes. Correction and retraction checks were current as of 2 August 2026 and do not guarantee that the record will remain unchanged.
Sources
- Primary article and version of record ↗
- Transparent peer-review file ↗
- BioStudies source data ↗
- Author code repository ↗
- PubMed record ↗
Current version 3. Corrected remaining mojibake punctuation and character encoding throughout the public evidence brief; no claims, citations, or metadata changed.
- Version 1 · 2026-08-03 — Initial publication after source verification and editorial review.
- Version 2 · 2026-08-07 — Corrected the DNA–RNA headline punctuation and repaired its character encoding; no other publication content or metadata changed.
- Version 3 · 2026-08-07 — Corrected remaining mojibake punctuation and character encoding throughout the public evidence brief; no claims, citations, or metadata changed.