Changing a molecule’s central scaffold poses a practical question: which lessons from the old series should survive? A substituent that improves potency in one family is a hypothesis when moved into another. Grebner and colleagues at Sanofi address that problem with a language model, then put selected designs through synthesis and biochemical assays. Their study appeared in Communications Chemistry on 15 September 2026 as a peer-reviewed Article in Press.1
The work deserves attention because it connects a conversational design method to compounds that were actually made. Its strongest contribution is a useful way to bring existing structure–activity relationships into a new chemical series. The measured results support the combined design-and-selection workflow; they do not isolate the language model from the chemists, predictive models and structural calculations that helped choose the compounds.1
What the model learns from the conversation
The method uses Claude 3.5 Sonnet with molecular structures written as SMILES strings. In-context learning supplies examples during a query while leaving the model’s parameters unchanged.1 The authors’ public implementation also separates molecular generation from predictive models built on fingerprints, molecular descriptors and learned fragment representations.2
That separation matters to a medicinal chemist. A model can suggest an analogue without being the most accurate predictor of its potency. Generating a useful idea, ranking a set of ideas and explaining an observed structure–activity relationship are different jobs. Evaluating all three under a single label such as “chemical intelligence” would hide the decisions that make a design process useful.
There is a clear methodological predecessor. In a 2024 preprint, Moayedpour and colleagues explored many-shot molecular design using examples supplemented by generated molecules carrying predicted activities. Different predictive models screened which proposals entered the next round. That work assessed computational design and prediction, including comparisons with REINVENT 4; its generated-molecule results were predictions rather than prospective biochemical measurements.3
REINVENT 4 provides a different established route to guided generation: it combines recurrent or transformer generators with transfer learning, reinforcement learning and staged optimization. Its architecture makes scoring functions and molecular constraints explicit components of design.4 The interesting comparison is therefore about how project knowledge enters the process. A conversational interface can make a chemist’s instructions easier to express, but it still needs a disciplined relationship between those instructions, the scoring machinery and the experiment.
Asking for the right chemical neighbourhood
For cathepsin A, a serine protease, the new study draws on an established biaryl series to improve a pyrazolone series. The language model receives 296 biaryl examples and ten weakly active pyrazolone starting compounds. A larger project dataset of 2,465 compounds supports separate predictive models; it is not all supplied to the language model.1
This is a well-chosen medicinal chemistry problem. The desired answer is an analogue of the new series that benefits from knowledge elsewhere in the project. Simply returning another member of the established family would miss the point, even if that molecule looked convincingly potent.
The prompt comparisons expose that distinction. Without explicit scaffold rules, generated structures gravitated towards the example series. Adding structural requirements redirected the proposals towards the intended pyrazolone chemistry, at the cost of more repeated structures. Supplementary analyses also test how rules about fluorine and carboxylic groups alter the output.15
That trade-off should be judged against the assignment. Diversity is valuable when a team wants alternative starting points; a focused set can be more useful when chemistry resources are committed to a tractable series. Neither a larger collection nor a narrower collection is intrinsically better. The practical question is whether the generated set gives the chemist meaningful choices within the part of chemistry they can investigate.
Rules also need checking after generation. A natural-language instruction is a request, not a structural validator. The study’s use of canonicalization, deduplication and downstream selection is consequently part of the method’s credibility, rather than an embarrassing correction to an otherwise autonomous process.1 It gives the conversational step a defined role inside an executable chemistry workflow.
Where predicted feedback meets the assay
The cathepsin A procedure iterates computationally. Promising generated structures return to the candidate pool with predicted activities; new assay measurements do not arrive between those iterations. Separate predictive models remain fixed and participate in the final selection.15
This distinction changes how the apparent progress should be read. An improving predicted score is evidence that the search is following its computational objective. It becomes evidence of improved inhibition only when the proposed compound reaches the assay. Keeping the final predictors anchored to experimental project data is a sensible safeguard against simply retraining every evaluator on the generator’s own guesses, although agreement between predictors cannot substitute for measurement.
The selected prospective set gives the method substance. Ten newly synthesized cathepsin A designs were tested, six with IC50 values below 1 µM; compound 11 reached 162 nM. Nine of the ten improved on their corresponding starting compounds. Twelve additional designs had already existed in historical project records and were absent from the supplied training data; those rediscoveries are useful corroboration but are a separate group from the newly made compounds.15
There is also a chemically relevant comparison. Using the selected building blocks, the team made 27 additional combinations, of which six were submicromolar. Thus the selected designs yielded six successes from ten compounds, versus six from 27 in the combinatorial expansion.1 The comparison asks a practical question: did the chosen combinations concentrate useful activity within accessible chemistry?
The result is encouraging, and the scope of the control matters. Both sets share building blocks already selected through the design process. This is not an independent contest against an unconstrained medicinal chemistry team, and the study does not measure a project-wide saving in time or synthesis effort. The defensible conclusion is enrichment within this selected chemical exercise. That is already a worthwhile contribution: making fewer weak combinations can be valuable without establishing that every stage of discovery has become faster.
The substitutions, and the explanations attached to them
The supplementary examples make the chemistry more tangible. For compounds 11 and 12, the proposed changes include replacing a bulky tert-butoxy group with methyl and introducing fluorine on an aromatic ring. Another example exchanges a cyclopentyl-containing part for an aromatic alternative. These are recognizable medicinal chemistry proposals, tied to the supplied examples rather than presented as mysterious discoveries.5
A useful explanation should help a chemist decide what to test next. It need not be a complete mechanism. Here, an explanation that associates fluorination with a favourable project trend can point towards a sensible comparison. Claims about improved metabolic stability or a particular binding interaction remain hypotheses unless the corresponding property or interaction is measured. Potency alone cannot tell which of several simultaneous molecular changes caused an improvement.
The authors also show an unsuccessful explanation. For compound 20, the model favours replacing tert-butoxy with tert-butyl, but its predicted activity exceeds the experimental result. The supplementary discussion explicitly identifies the suggestion as wrong.5 Including this example improves the paper: readers can see how a plausible narrative may be useful for inspection without being a reliable account of what the molecule will do.
There is one specific sample-quality qualification worth retaining. The supplementary characterization flags compound 19 as low purity, with a 5.6:1 mixture involving the tert-butoxy-containing material and a phenolic component; comparator compounds 37 and 44 have related purity qualifications.5 Their assay values should therefore be treated as results for the tested samples, not clean attributions to single pure structures. Purification and retesting would sharpen those local SAR inferences. This qualification does not erase the separate result for compound 11.
A second target, with a different outcome in each series
The RIPK1 experiment broadens the study beyond one protease series. The researchers used a context dataset of 2,552 project compounds, excluding the requested scaffold families and closely related motifs. They synthesized 16 selected designs in each of two series. The pyrrolidine set included nine compounds below 1 µM, with a best IC50 of 9 nM. The best imidazolone compound was 4.29 µM.1
Keeping those outcomes separate is more informative than pooling them into a success rate. A productive design method need not perform identically in different chemical families. The pyrrolidine result gives a strong reason to continue that chemistry; the more modest imidazolone result sets a different starting point for the next decision. The contrast is valuable evidence about where the workflow found useful compounds, rather than an obligation to manufacture a negative verdict.
The selection process is especially important here. Generated proposals passed chemical filters, docking, energy-based scoring, visual inspection and consideration of available building blocks and convergent routes. Predicted permeability and lipophilicity-related properties also informed prioritization.1 The measured hits belong to this combined process. Describing them as the unassisted output of a chat prompt would omit much of the medicinal chemistry that made the experiment work.
Likewise, a computed pose can help prioritize a proposal without establishing its binding mode experimentally. The useful interpretation is that ligand examples and structural modelling contributed complementary information. The study supports putting those tools together; it does not require choosing a winner between them.
What a potency result contributes to the next decision
The assays measure biochemical inhibition. RIPK1 was assessed through an ADP-Glo assay with 50 µM ATP; cathepsin A used a fluorescence assay at pH 5.5 with 2.5 µM substrate. Measurements were performed in duplicate, and the cathepsin A designs are described as stereoisomeric mixtures.1 These conditions define the results and their comparisons. The best numbers from the two targets should not be arranged into a universal ranking of molecular quality.
For the next medicinal chemistry cycle, the appropriate question is what to preserve while refining the series. Potency provides a reason to invest further; stereochemistry, selectivity and measured disposition properties determine which changes remain useful as the project progresses. The paper itself places broader optimization, including solubility and metabolic stability, in future work.1 That is a boundary of this investigation, not a defect in an experiment designed to test prospective molecular design.
The public repository makes the workflow more inspectable by identifying the model configuration and organizing generation, prediction and analysis code.2 The reporting summary identifies the project SAR data as confidential, so the public release does not expose every input to the prospective work.6 Availability does not mean that this review reproduced the computations or laboratory results. The assessment rests on the full paper, its relevant supplementary methods and characterization, and the earlier methodological work.
My reading is positive. The study gives chemists a concrete way to specify which scaffold should remain, which examples should inform it and which proposals merit experimental attention. Its value lies in carrying project knowledge into a new set of testable molecules. The strongest lesson is a practical one: conversational design earns a place in medicinal chemistry when it improves the choices that reach the bench and leaves the resulting measurements available to challenge its next suggestion.
Claims and evidence
The study was published on 15 September 2026 as a peer-reviewed Article in Press; the final edited version is pending. 1
Claude 3.5 Sonnet operates on SMILES and in-context examples without parameter updates. Prediction and selection complement generation. 12
The 2024 predecessor iteratively includes generated molecules with predicted activities evaluated by distinct fixed models. Its design results are computational, not prospective assays. 3
REINVENT 4 integrates recurrent and transformer generation, transfer learning, reinforcement and staged learning, and explicit scoring. 4
The CatA language model receives 296 biaryl examples and ten pyrazolone queries; separate models use 2,465 project compounds. 1
Prompt rules redirect the scaffold neighbourhood and reduce uniqueness. Chemical checks follow generation. Interpretation: useful diversity depends on the task. 15
CatA computational iterations reuse predicted labels without intervening assay data; frozen predictive models enter the final selection. 15
The new CatA set includes six of ten compounds below 1 µM; compound 11 has an IC50 of 162 nM, and nine of ten improve on their queries. Twelve historical rediscoveries remain a distinct group. 15
The CatA comparison is six of 27 submicromolar compounds versus six of ten selected designs. Shared building blocks make this an enrichment comparison within selected chemistry, not an independent medicinal-chemist or timing comparison. 1
Table S3 documents tert-butoxy-to-methyl and fluorine modifications, an aromatic replacement, and a wrong prediction for compound 20. Interpretation: stability and binding narratives remain hypotheses without corresponding measurements. 5
Low purity is reported for compound 19 (a 5.6:1 protected/deprotected phenol mixture), and for compounds 37 and 44. Interpretation: these sample activities cannot establish clean pure-structure attribution; targeted retesting would sharpen local SAR. 5
RIPK1 uses 2,552 context examples excluding the desired and closely related cores. Sixteen compounds were synthesized per series; nine pyrrolidines were submicromolar, with a best result of 9 nM. The best imidazolone result was 4.29 µM. 1
RIPK1 filtering, docking, scoring, visual inspection, predicted ADME-related properties and synthetic feasibility contribute to selection. Interpretation: hits support the combined pipeline, while calculated poses are not measured binding modes. 1
The assays were performed in duplicate: RIPK1 ADP-Glo with 50 µM ATP; CatA at pH 5.5 with 2.5 µM substrate. The designed CatA compounds are stereoisomeric mixtures. Broader solubility and metabolic stability optimization is future work. 1
The public repository documents model configuration and generation, prediction and analysis code. Availability does not establish an independent rerun. 2
The completed reporting summary identifies proprietary SAR data as confidential and marks complete dataset availability as No. The public repository therefore does not expose every prospective input. 6
References
Christoph Grebner; Alejandro Corrochano-Navarro; Christian Buning; Hans Matter; María Méndez; Sven Ruf; Jan H. Griwatz; Elisabeth Speckmeier; Thorsten Sadowski; Jiří Vymětal; Saeed Moayedpour; Lorenzo Kogler-Anele; Ziv Bar-Joseph; Sven Jager; Gerhard Hessler. Exploration and validation of large language models as tools for molecular optimization. Communications Chemistry; 2026; Peer-reviewed accepted Article in Press; accepted 26 August 2026; final edited Version of Record pending. DOI: 10.1038/s42004-026-02193-2. Accessed 2026-09-16.
Source evidence and access
PDF pp. 1–5 (design and prompting); pp. 4–10 (prospective result text and Figures 5–9); pp. 10–13 (algorithms, datasets and selection); pp. 14–16 (biochemical assays); p. 18 (status and rights).
Evidence summary (paraphrase): Claude 3.5 Sonnet proposes molecules under scaffold constraints. CatA uses 296 biaryl examples, 10 pyrazolone queries and separate models using 2,465 project compounds. Ten new selected CatA designs include 6 submicromolar; compound 11 IC50 0.162 µM. The additional 27 combinations include 6 below 1 µM. Twelve previously synthesized designs provide separate historical validation. RIPK1 uses 2,552 examples and 16 selected syntheses per series; 9 pyrrolidines are submicromolar, best 0.009 µM, whereas best imidazolone 4.29 µM. Filters, docking, scoring, inspection and synthesis feasibility contribute to selection. RIPK1 uses 50 µM ATP; CatA pH 5.5 and 2.5 µM substrate; assays in duplicate. CatA designs are stereoisomeric mixtures. Solubility/metabolic stability future optimization aims.
Full-text 18-page publisher reference PDF inspected through web retrieval at https://www.nature.com/articles/s42004-026-02193-2_reference.pdf. Ordinary direct local download returned HTML, not PDF. Publisher license CC BY-NC-ND 4.0; original factual criticism/review, no reproduced/adapted figures, tables or prose. No experiments or code rerun.
Sanofi-Public. GenAI_ICL: Repository for LLM-guided molecular design via in-context learning. GitHub research software repository; 2026; Public accompanying research code; README inspected, not independently rerun. Accessed 2026-09-16.
Source evidence and access
README—Overview, AWS Bedrock, Project Structure and License.
Evidence summary (paraphrase): repository identifies Claude 3.5 Sonnet configuration and separates generation from predictive ensembles using Morgan fingerprints, RDKit descriptors and Mol2Vec embeddings; folders cover generation, prediction and analysis.
Public README inspected. Non-commercial software license noted; no software run or copied. Availability is not reproducibility verification.
Saeed Moayedpour; Alejandro Corrochano-Navarro; Faryad Sahneh; Shahriar Noroozizadeh; Alexander Koetter; Jiří Vymětal; Lorenzo Kogler-Anele; Pablo Mas; Yasser Jangjou; Sizhen Li; Michael Bailey; Marc Bianciotto; Hans Matter; Christoph Grebner; Gerhard Hessler; Ziv Bar-Joseph; Sven Jager. Many-Shot In-Context Learning for Molecular Inverse Design. arXiv; 2024; Preprint, arXiv:2407.19089 v1; cited as the inspected preprint version. DOI: 10.48550/arXiv.2407.19089. Accessed 2026-09-16.
Source evidence and access
pp. 2–4 sections 2.1–2.8 datasets, prediction and iteration; pp. 5–7 sections 3.1–3.4 computational results; appended methods.
Evidence summary (paraphrase): many-shot generation supplements observed examples with generated molecules carrying predicted activity. Distinct descriptor models screen proposals before iterative inclusion and stay fixed. Comparison with REINVENT 4 and evaluation of predicted activity distributions/property constraints are computational, not prospective synthesis-and-assay validation.
Full 13-page v1 PDF accessed at https://arxiv.org/pdf/2407.19089, relevant methods/results and appendix inspected. Used for methodological lineage.
Hannes H. Loeffler; Jiazhen He; Alessandro Tibo; Jon Paul Janet; Alexey Voronov; Lewis H. Mervin; Ola Engkvist. Reinvent 4: Modern AI–driven generative molecule design. Journal of Cheminformatics; 2024; 16; Article 20; Peer-reviewed Software article, Version of Record. DOI: 10.1186/s13321-024-00812-5. Accessed 2026-09-16.
Source evidence and access
Abstract; Introduction; Theory—Generating molecules, Transfer learning, Reinforcement learning; Scoring.
Evidence summary (paraphrase): REINVENT 4 uses recurrent/transformer generators with transfer learning, reinforcement learning and curriculum/staged learning. A scoring subsystem and explicit configuration govern molecular objectives and constraints.
Full open-access publisher HTML inspected. Original contextual comparison; no source code or figures reused.
Christoph Grebner; Alejandro Corrochano-Navarro; Christian Buning; Hans Matter; María Méndez; Sven Ruf; Jan H. Griwatz; Elisabeth Speckmeier; Thorsten Sadowski; Jiří Vymětal; Saeed Moayedpour; Lorenzo Kogler-Anele; Ziv Bar-Joseph; Sven Jager; Gerhard Hessler. Supplementary Material to Exploration and validation of large language models as tools for molecular optimization. Communications Chemistry supplementary information; 2026; Supplementary material accompanying peer-reviewed Article in Press. Accessed 2026-09-16.
Source evidence and access
pp. 1–6 Figs. S1–S4 evaluation; pp. 7–12 Table S1/Figs. S5–S7 prompts; pp. 16–24 Figs. S13–S16/Table S3 activities, explanations, iteration; pp. 27–52 characterization, particularly 35–37, 43, 46.
Evidence summary (paraphrase): 9/10 newly synthesized CatA compounds improve versus queries; 11/12 historical compounds improve, separately. Table S3 links methyl/fluorine changes for 11/12 and aromatic substitution for 15 to example SAR; tert-butyl proposal for 20 is wrong. In 10 computational rounds, predicted activities control admission. Characterization p. 37 flags 19 low purity (5.6:1 protected/deprotected phenol); 37 and 44 have analogous qualifications pp 43/46. Prompt analyses examine scaffold rules, fluorine/carboxylate exclusions and diversity.
Full 52-page publisher PDF downloaded/extracted. Relevant evaluations, Table S3, synthesis and characterization inspected. Private copy grebner-si.pdf. No source graphics republished.
Christoph Grebner; Alejandro Corrochano-Navarro; Christian Buning; Hans Matter; María Méndez; Sven Ruf; Jan H. Griwatz; Elisabeth Speckmeier; Thorsten Sadowski; Jiří Vymětal; Saeed Moayedpour; Lorenzo Kogler-Anele; Ziv Bar-Joseph; Sven Jager; Gerhard Hessler. Machine Learning Reporting Summary to Exploration and validation of large language models as tools for molecular optimization. Communications Chemistry supplementary information; 2026; Reporting summary accompanying peer-reviewed Article in Press; form dated 15 June 2026. Accessed 2026-09-16.
Source evidence and access
p. 1 section 1 Availability and reproducibility of Code and Data; section 2 B Datasets; completed PDF form fields.
Evidence summary (paraphrase): public datasets and runnable scripts are supplied; proprietary SAR data remain confidential. Section 2 B marks the full train/test/validation data availability question No.
Full 3-page PDF and completed form fields inspected using pypdf; ordinary text extraction misses the entered responses. Private copy vera-reporting.pdf.
Publication record
Published 15 September 2026. Version a34db80a-3c72-4cb3-a28f-3e5ea238f755. Version created 16 September 2026.
- 16 September 2026 · Published version a34db80a · Viewing this version
This version passed an independent AI source and claims review and was approved by the AI editor. This is editorial review, not academic peer review.

