Research papers
1. MIRAGE tests affinity prediction on unfamiliar protein families
This unreviewed benchmark finds that Nesso-1 affinity accuracy increases with historical protein-family representation in public structural data. A random-forest baseline trained without the test family performs better on sparsely represented families.1
Why it matters: Breaking results down by family familiarity can reveal weaknesses hidden by an overall benchmark score.1
Limits: Public family counts are a proxy, not the models’ disclosed training histories. The study cannot establish training leakage or within-target compound ranking; Boltz-2 comparisons have limited coverage and the external test covers one target.1
2. Designing molecules around more than one conformation
An unreviewed preprint introduces a generator guided by multiple molecular shapes, interaction features or protein pockets. In computational case studies, combined conditions improved dual-target docking scores and preference for an active receptor state over single-state conditioning.2
Why it matters: The method lets researchers specify both wanted and unwanted states within one generation framework.2
Limits: Docking does not establish binding or agonism. Ensemble sampling is approximate, and property control is directional rather than calibrated.2
3. HiFi-Mol combines molecular graphs with learned fingerprints
HiFi-Mol learns separately from molecular graphs and fingerprint descriptors, then combines those representations for property prediction. On eight MoleculeNet classification datasets split by molecular scaffold, the combined model had the highest average ROC-AUC among the reported baselines.3
Why it matters: The results support learning complementary chemical representations when labeled property data are limited.3
Limits: This is retrospective benchmark evidence; gains vary by dataset and combining the views does not consistently improve regression. The arXiv manuscript reports CIKM 2026 acceptance; a final proceedings version was not checked.3
4. SoupFold combines information inside co-folding models
An unreviewed preprint combines internal representations from different structure-prediction models while keeping their original weights fixed. On 517 shared FoldBench protein–ligand interfaces, ESMFold2 with SoupFold reached 65.76% structural success, compared with 63.46% without it, using confidence-selected predictions.4
Why it matters: The experiment suggests models can contribute useful structural information beyond simply choosing among their finished predictions.4
Limits: Transfer mappings are fitted on representations of the target complexes; this is not a test of fixed mappings on unseen targets. Evaluation retains cases processed by all four models with matching token correspondence and measures structure accuracy, not binding affinity.4
Research news
1. Pharma consortium reports gains from federated co-folding
Apheris reports that five pharmaceutical partners jointly fine-tuned OpenFold3 Preview 2 while retaining their structural data locally. On 1,056 held-out private complexes, the fraction meeting its high-quality interface threshold rose from 35.6% to 52.1%.5
Why it matters: The company-reported comparison tests whether distributed proprietary structures can improve a shared protein–ligand prediction model.5
Limits: This is a company technical announcement, not a peer-reviewed paper. Data and weights remain private; evaluation used participating partners' held-out projects, not a fully external test, and did not establish binding affinity.5
References
Mehdi Yazdani-Jahromi; Sanjay Padhi; Ivan Garibay. MIRAGE: Measuring Interpolation and Redundancy in Affinity GEneralization. arXiv; 2026; Preprint, arXiv v1; unreviewed. DOI: 10.48550/arXiv.2609.14491. Accessed 2026-09-15.
Source evidence and access
arXiv v1 submitted 13 September 2026 12:56:32 UTC. HTML sections 2–3, Table 1; Appendix O limitations and Table 24.
Evidence paraphrase: family-disjoint controls and matched strata test public-family-support dependence. Novel-family random forest comparison against Nesso-1 excludes zero. Exact training datasets cannot be reconstructed and within-target ranking is not estimable.
Full HTML benchmark methods, principal results and extended limitations read; code not rerun.
Ross Irwin; Alessandro Tibo; Jon Paul Janet; Simon Olsson. Ensemble-Conditioned Molecular Design. arXiv; 2026; Preprint, arXiv v1; unreviewed. DOI: 10.48550/arXiv.2609.15077. Accessed 2026-09-15.
Source evidence and access
arXiv v1 submission history (14 Sep 2026 05:49:37 UTC); full HTML sections 3, 4.3 and 5 Limitations.
Evidence paraphrase: multiple conditions are composed during generation. Section 4.3 reports docking-based dual-target and receptor-state comparisons. Section 5 explicitly restricts claims because docking and approximate conformer sampling cannot establish experimental success.
Full HTML methods, case studies and limitations read; code and experiments not rerun.
Gwang-Hyeon Yun; Jong-Hoon Park; Bing Hu; Helen Chen; Anita Layton; Young-Rae Cho. Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints. arXiv; 2026; arXiv v1 manuscript; authors report acceptance as a Full Research Paper at CIKM 2026; proceedings version not verified. DOI: 10.48550/arXiv.2609.15611. Accessed 2026-09-15.
Source evidence and access
arXiv v1 submission history (14 Sep 2026 14:07:19 UTC) and Comments; full HTML sections 4–6, Tables 2–3.
Evidence paraphrase: Table 2 puts combined HiFi-Mol highest on average across eight scaffold-split classification datasets. Table 3 and the conclusion show mixed regression gains. Section 4 describes separate graph and fingerprint encoders.
Full HTML methods, evaluation tables and conclusions read; no independent model rerun.
Hyosoon Jang; Taewon Kim; Sungsoo Ahn. Synthesizing State-of-the-Art Structure Predictions from Soup of Co-folding Models. arXiv; 2026; Preprint, arXiv v1; unreviewed. DOI: 10.48550/arXiv.2609.15552. Accessed 2026-09-15.
Source evidence and access
arXiv v1 submitted 14 September 2026 13:36:25 UTC. PDF page 1 author list; section 2/Table 1; section 3; Appendix A.1.
Evidence paraphrase: Table 1 compares confidence-selected best-of-five results. Success requires ligand RMSD below 2 angstrom and lDDT-PLI above 0.8. Appendix A.1 fits transfer networks on target representations and retains exact token correspondence across all four models.
Full HTML results and methods read; PDF title page and Table 1 also checked. No independent rerun.
Apheris; AI Structural Biology Network. Federated Training Dramatically Improves the Accuracy of Protein-Ligand Co-folding on Private Pharma Structures. Apheris technical announcement; 2026; Company technical announcement; not presented as a peer-reviewed paper. Accessed 2026-09-15.
Source evidence and access
Published on date; Main result; Evaluation setup; How the AISB-1-Fed model was trained; Interpretation and limitations.
Evidence paraphrase: Main result gives PL-lDDT ≥0.8 fractions of 35.6% and 52.1%. Evaluation setup specifies 1,056 shared-success cases and project-level splitting. Interpretation and limitations identifies partner-held-out scope and no affinity evaluation.
Full public technical announcement read; underlying private structures and trained model weights unavailable.
Publication record
Published 2026-09-15.
Sources, selection and claims were checked in an independent AI editorial review, followed by the AI editor's approval. This is not academic peer review.
- 2026-09-15 · Version 73c21269 · Viewing this version