1. A Chemical Foundation Model Benchmark for PFAS Bioactivity Prediction: Multi-Task Fine-Tuning Beats Fingerprint Baselines, PFAS Domain Pretraining Does Not Help1
A PFAS bioactivity benchmark shows how representation and evaluation choices can change apparent model advantages. Across 156 ToxCast endpoints, multitask MolFormer outperformed tested classical representations under cluster-based splits, while PFAS-specific pretraining did not help. The analysis also identifies weaknesses in fingerprint baselines for repetitive fluorinated structures, offering a more informative comparison for molecular screening. Our account is restricted to the deposited abstract.1
2. PFArena: Benchmarking Language Models for Protein Modification2
PFArena compares six protein language models, six general language models and five agent systems across four protein-modification interfaces. The unreviewed benchmark finds that relative performance changes with available target-specific fitness data: protein models favour open-ended mutation generation, while language models and agents perform strongly in informed ranking. This helps match tools to actual discovery decisions. Increasing mutation depth remains difficult; this account is abstract-only.2
3. MEGPNM as a multiscale edge-aware GAT network with hybrid pooling predicting permeability of non-peptidic macrocycles3
MEGPNM predicts artificial-membrane permeability of non-peptidic macrocycles using a graph network that combines local bond information with representations at several scales. It outperformed 17 comparators across random, activity-cliff and scaffold splits, and was tested on 187 external molecules. The study offers a screening tool for this difficult chemical class, while its two-dimensional representation and assay endpoint limit conclusions about conformational effects or oral absorption.3
4. A Congener-Substitution Strategy for Docking Heavy Metal Protein-Ligand Complexes with AutoDock Vina4
An unreviewed docking study substitutes first-row metal congeners for heavier centres so that metal-containing complexes can be processed with AutoDock Vina. Tests against roughly 60 crystal structures recovered rigid docking poses in 66% of cases, with a smaller flexible-ligand evaluation showing partial success. The workaround broadens accessible screening, but it does not replace dedicated treatment of metal chemistry. Our account is restricted to the abstract.4
5. ProteoEM: probabilistic protein abundance estimation from iterative affinity traces5
ProteoEM apportions ambiguous affinity traces among candidate proteoforms instead of assigning each trace a single identity. In simulations, this probabilistic treatment recovered mixture composition more accurately than hard calls and grouped indistinguishable candidates. The unreviewed method could improve quantitative interpretation of single-molecule assays, but missing reference proteoforms and informative missing data introduced bias; experimental validation is still needed. This account is abstract-only.5
References
Yue Zhu; Qingyang Liu. A Chemical Foundation Model Benchmark for PFAS Bioactivity Prediction: Multi-Task Fine-Tuning Beats Fingerprint Baselines, PFAS Domain Pretraining Does Not Help. Journal of Chemical Information and Modeling; 2026; Peer-reviewed journal article. DOI: 10.1021/acs.jcim.6c02202. Accessed 2026-09-26T16:20:08.087Z.
Source evidence and access
Deposited Abstract and first-online/posting metadata: https://api.crossref.org/works/10.1021/acs.jcim.6c02202
The deposited abstract describes 156 ToxCast endpoints, scaffold degeneracy affecting 89% of compounds, cluster splits, MolFormer AUROC 0.764, eight classical representations with best AUROC 0.688, and no gain from PFAS domain pretraining.
Account restricted to the original publisher-deposited abstract and metadata; full text not inspected.
Yawen Ouyang; Xinbo Zhang; Ziyuan Ma; Yixin Wu; Wenbin Liao; Feiran Zhang; Wenjie Li; Lihao Wang; Hao Wang; Xiaoqing Zheng; Xuefeng Yan; Lei Bai; Ya-Qin Zhang; Shuyi Zhang; Wei-Ying Ma; Dahua Lin; Bowen Zhou; Hao Zhou. PFArena: Benchmarking Language Models for Protein Modification. arXiv; 2026; Preprint v1; not peer reviewed. DOI: 10.48550/arXiv.2609.28921. Accessed 2026-09-26T16:08:46.332Z.
Source evidence and access
Original abstract and v1 submission history at https://arxiv.org/abs/2609.28921
The original abstract specifies six PLMs, six LLMs and five agents, four controlled task interfaces, single-mutant generation versus multi-mutant ranking, changing performance with fitness-data access, and difficulty at increased mutation depth/search-space size.
Original arXiv abstract and submission history inspected; account is abstract-only.
Shida He; Lesong Wei; Weidong Ye; Quan Zou; Feng Zhang. MEGPNM as a multiscale edge-aware GAT network with hybrid pooling predicting permeability of non-peptidic macrocycles. Communications Chemistry; 2026; 9; Article 301; Peer-reviewed journal article. DOI: 10.1038/s42004-026-02194-1. Accessed 2026-09-26T16:20:08.087Z.
Source evidence and access
Abstract; Results: performance comparison and external validation; Methods: Datasets; Conclusions.
Methods curate 3,412 unique NPMMPD macrocycles from PAMPA records. Results compare 17 models over random/cliff/scaffold splits and 187 external compounds from ten sources (MEGPNM RMSE 0.750). Conclusions identify missing explicit 3D conformational information; the measured endpoint is artificial-membrane permeability.
Open-access original article inspected: abstract, datasets, split comparisons, external validation and conclusions.
Miroslava Nedyalkova; Simone Bernardotto; Fabio Zobi. A Congener-Substitution Strategy for Docking Heavy Metal Protein-Ligand Complexes with AutoDock Vina. ChemRxiv; 2026; Preprint v1; not peer reviewed. DOI: 10.26434/chemrxiv.15009469/v1. Accessed 2026-09-26T16:20:08.087Z.
Source evidence and access
Deposited Abstract and first-online/posting metadata: https://api.crossref.org/works/10.26434/chemrxiv.15009469/v1
The deposited abstract reports first-row congener substitution preserving coordination geometry, approximately 60 crystal complexes with 66% rigid docking success, and 18 biotin-conjugate complexes with 60% full or partial pose recovery in flexible docking.
Account restricted to the original publisher-deposited abstract and metadata; full text not inspected.
Narayanan Raghupathy. ProteoEM: probabilistic protein abundance estimation from iterative affinity traces. arXiv; 2026; Preprint v1; not peer reviewed. DOI: 10.48550/arXiv.2609.29155. Accessed 2026-09-26T16:08:46.348Z.
Source evidence and access
Original abstract and v1 submission history at https://arxiv.org/abs/2609.29155
The original abstract describes expectation-maximisation using independently calibrated probe responses, simulated recovery versus hard calls, indistinguishable groups, and bias from informative missingness/absent reference proteoforms; it calls for experimental molecule-level validation.
Original arXiv abstract and submission history inspected; account is abstract-only.
Publication record
Published 2026-09-26.
Sources, selection and claims were checked in an independent AI editorial review, followed by the AI editor's approval. This is not academic peer review.
- 2026-09-26 · Version 8e572384 · Viewing this version