Imagine stripping a modelling problem down until only a deliberately simple representation remains. If that model still performs well, the result asks something of the benchmark as well as of the model. Molecular set representation learning, published in 2024, makes this a useful question for chemistry.1
This is a launch methods retrospective, not a claim of newly released research. We revisit it because choosing how to represent a molecule remains inseparable from choosing what an evaluation can establish.
Keep the atoms, relax the explicit wiring
Boulougouri and colleagues represent molecules as multisets of atom-feature vectors; the simplest model does not encode explicit graph connectivity, although its atom invariants can retain local structural information.1 Multiset means that repeated entries still count. The order of the atoms should not determine the prediction.1
That qualification prevents an easy misunderstanding. This is not equivalent to handing a model nothing but a molecular formula. The features supplied to each atom are part of the input information. We would therefore compare representations by listing what they know, rather than by judging their apparent simplicity from a diagram.
Let a simple baseline challenge the task
On the paper's MoleculeNet comparisons, the simplest set model performs competitively with graph-based alternatives; the authors discuss benchmark limitations as one possible explanation.1 Their argument does not establish that molecular connectivity is generally irrelevant. It raises a question about which distinctions the tested datasets force a model to learn.
For a reader choosing an evaluation, our suggested exercise is to state which chemical change ought to matter for the task, then ask whether the benchmark distinguishes it. A high score from a sparse representation can be valuable evidence about the dataset even when that representation would be insufficient for a more demanding scientific question.
The purpose of this exercise is not to dismiss a benchmark after one surprising result. It is to use the surprise to design a more discriminating comparison. The model and the test should be investigated together.
Sets can also supplement graphs
The study also replaces a graph model's usual pooling operation with a set representation layer, showing improvements in reported comparisons.1 Thus the proposal includes a way to combine graph-derived information with set-based aggregation, rather than asking every application to choose one representation permanently.1
Our interpretation is that this makes the paper a useful prompt for controlled ablations: change one component, preserve the rest and inspect which tasks improve or deteriorate. A well-described negative result would be as helpful as another headline average, because it would reveal where the component stops earning its complexity.
Match the claim to the test
AiChemEx has not rerun these models. We would carry forward a reporting rule: show the simple baseline, describe its exact input features and avoid treating benchmark parity as a universal statement about chemistry.
The practical question is what a representation must preserve to answer a particular question well. A set of atoms can help expose an evaluation's assumptions; it does not make those assumptions disappear. That is why a modest baseline can sometimes be one of the most revealing parts of a molecular modelling paper.
What this does not establish
- The simple representation still contains atom invariants and is not merely a molecular formula.
- Benchmark performance does not establish that connectivity is unimportant for chemistry or all prediction tasks.
- This is a 2024 study revisited for launch; code, data and performance comparisons were not independently rerun.
Claims and evidence
References
Maria Boulougouri, Pierre Vandergheynst and Daniel Probst. Molecular set representation learning. Nature Machine Intelligence; 2024; 6; 754–763; peer-reviewed journal article. DOI: 10.1038/s42256-024-00856-0. Accessed 2026-09-15.
Source evidence and access
Main, multiset representation definition; Results and discussion, MoleculeNet comparisons and set-enhanced graph model
it does not encode any explicit connectivity of the molecular graph
Publisher full-text HTML: cited sections and bibliographic metadata inspected; code and supplementary analyses not independently reproduced.
Publication record
Published 15 September 2026. Version 44c25fbe-6107-4c6d-b467-2de1179e9f26. Version created 15 September 2026.
- 15 September 2026 · Published version 44c25fbe · Viewing this version
- 15 September 2026 · Published version 9760e667
This version passed an independent AI source and claims review and was approved by the AI editor. This is editorial review, not academic peer review.

