Imagine receiving a list of proposed molecules accompanied by elegant explanations. The immediate editorial question is what evidence belongs beside each structure. A drawing, a predicted property and a measured result should remain distinguishable all the way through a discovery story.
A prompt can specify the search
SmileyLlama, reported in May 2026, adapts a general language model to produce molecular representations through supervised fine-tuning. Direct preference optimization then improves its response to requested properties; the study also explores docking-oriented optimization. Its authors identify a trade-off: stronger adherence can reduce diversity. They also report difficulties in data-poor settings such as macrocycle generation. These results concern a molecular-generation method, and do not establish that its suggestions are safe or effective medicines. 1
For our coverage, a useful generation demonstration should disclose its rejected outputs as well as its selected illustrations. We would want to know whether the proposed collection remains varied after the selection criteria are applied. That is a reporting preference, rather than a further claim about this model's measured performance.
Another route starts with chemical data
ChemFM offers a different example: a three-billion-parameter model pretrained on 178 million molecules and adapted to prediction, generation and reaction tasks. Its authors investigate dataset informativeness and scaling, including a comparison between UniChem and a larger molecular collection. The paper first appeared online in December 2025, despite belonging to a 2026 journal volume; an affiliation correction followed in February. Its benchmark programme is separate from SmileyLlama's, so the papers do not establish a direct winner. 2
A prediction needs a boundary
A third study, published in April, combines property prediction with molecular reconstruction. Its reconstruction-based unfamiliarity measure helps assess behaviour outside the training distribution. The team also reports wet-laboratory screening against two kinases, identifying seven compounds with low-micromolar potency and limited structural similarity to training molecules. Those experimental hits are a different category of evidence from a generated structure or docking score. They are not clinical validation. 3
Our interpretation
The useful comparison is about questions. What does a user ask the generator to optimise? What chemical examples shaped its representations? What warns the researcher that a prediction is being stretched? And which proposed candidates ultimately receive a measurement?
We would keep those answers in a candidate-by-candidate record, including failed measurements and disagreements. This article does not show that combining these three approaches improves discovery: they have not been tested as one system here. It sets out the evidence we would request before describing a convenient molecular interface as a dependable discovery workflow.
What this does not establish
- The three methods were studied separately; this article proposes no validated combined workflow.
- Docking and generated structures are computational outputs, not clinical efficacy or safety evidence.
- ChemFM first appeared online in December 2025; its 2026 volume number is not the original publication date.
Claims and evidence
SmileyLlama adapts a general LLM with supervised fine-tuning and preference optimization. 1
The study explores docking-oriented optimization; prompt adherence can narrow diversity, and macrocycles remain difficult. 1
ChemFM has a three-billion-parameter variant pretrained on 178 million molecules. 2
Tasks cover prediction, generation and reactions; scaling work compares UniChem with a larger collection. 2
Online publication preceded the 2026 volume; an affiliation correction followed in February. 2
Unfamiliarity combines reconstruction and property prediction to assess out-of-distribution behaviour. 3
Two kinase assays identify seven low-micromolar compounds with limited training-set similarity. 3
Online publication: 2026-05-11. 1
Online publication: 2025-12-18. 2
Online publication: 2026-04-22. 3
References
Joseph M. Cavanagh, Kunyang Sun, Andrew Gritsevskiy et al.. SmileyLlama: modifying large language models for directed chemical space exploration. 2026; peer-reviewed journal article. DOI: 10.1038/s43588-026-00986-y. Accessed 2026-09-15.
Source evidence and access
Discussion, paragraph beginning Even so, while DPO; Abstract; SmileyLlama with iMiner reinforcement learning
DPO improves adherence to the prompt, it does so at the cost of narrowing the distribution of properties or diversity
publisher full-text HTML sections inspected; supplementary data and code not independently reproduced
Feiyang Cai, Katelin Zacour, Tianyu Zhu et al.. ChemFM as a scaling law guided foundation model pre-trained on informative chemicals. 2025; peer-reviewed journal article. DOI: 10.1038/s42004-025-01793-8. Accessed 2026-09-15.
Source evidence and access
Abstract; Introduction, UniChem comparison; Change history 03 February 2026
ChemFM comprises 3 billion parameters and is pre-trained on 178 million molecules
publisher full-text HTML sections inspected; supplementary data and code not independently reproduced
Derek van Tilborg, Luke Rossen and Francesca Grisoni. Molecular deep learning at the edge of chemical space. 2026; peer-reviewed journal article. DOI: 10.1038/s42256-026-01216-w. Accessed 2026-09-15.
Source evidence and access
Abstract, final two sentences; Main, applicability-domain discussion
discovering seven compounds with low micromolar potency and limited similarity to training molecules
publisher full-text HTML sections inspected; supplementary data and code not independently reproduced
Publication record
Published 15 September 2026. Version 8d16b5e4-d093-47d5-96b4-5d0bbb6618b2. Version created 15 September 2026.
- 15 September 2026 · Published version 8d16b5e4 · Viewing this version
This version passed an independent AI source and claims review and was approved by the AI editor. This is editorial review, not academic peer review.