Agent-TIS: AI agents for PyRETIS path sampling¶
Agent-TIS is the PyRETIS integration for AI agents: large language model (LLM) agents and coding agents, such as Claude Code, that set up, run and analyse path-sampling simulations (TIS, RETIS, REPPTIS, permeability, infinite swapping) or develop and validate the code. It captures PyRETIS-specific procedural knowledge: method selection, move choice, interface placement, engine validation, and the primary literature that supports those choices. It helps an agent navigate the documentation and execute defined workflows; it does not replace convergence assessment, uncertainty analysis, or scientific judgement by the researcher.
Agent-TIS provides two agent skills, a user skill and a developer skill,
together with a compact, machine-readable documentation index,
llms.txt. The skills live in the
repository. A
skill is a folder of agent-agnostic Markdown files: a SKILL.md that
holds the rules, the workflow and a table of topic files, and the topic
files, which the agent reads when a step of the work needs them. The folder
is installed in an agent’s skills directory, and SKILL.md is loaded
when its task description applies. The .claude/skills/<name>/ path is
one installation example; the same files can guide any LLM agent that can
read local files and use the required tools.
The user skill — run and analyse simulations¶
skills/pyretis-pathsampling/ guides the configuration, execution,
and interpretation of path-sampling calculations. Its SKILL.md holds
the rules the agent always follows and a fixed minimal recipe (decision
table, a base input per engine, the keys to change, exact commands) that a
small or locally hosted model follows step by step. A capable model follows
its workflow: selection of a sampling task (TIS / RETIS / REPPTIS /
permeability), MD-engine configuration (internal/TurtleMD, GROMACS, LAMMPS,
CP2K, or OpenMM), definition of the order parameter and interface ladder,
initial-path preparation, execution with pyretis run, and analysis of
crossing probabilities and rates with pyretis analyse, reading the
topic file of each step:
File |
Read when |
|---|---|
|
choosing the task, the moves and their controls, high acceptance, priority shooting, infinite swapping or permeability |
|
defining the order parameter, the states A and B, or extra collective variables |
|
placing or re-spacing the interfaces, by hand or with
|
|
configuring an MD engine |
|
getting initial paths, running, monitoring, restarting, using several workers |
|
reading the report, WHAM, the REPPTIS MSM, skipping cycles |
|
a result looks wrong or a run stops, e.g. a local crossing probability of zero |
|
deciding whether to trust a setup, a rate or a speed-up: a known case with an analytical rate, invariance checks, efficiency |
|
citing the method behind a choice |
|
notes for a specific system |
The skill makes the following method choices explicit:
RETIS includes the initial-flux contribution and is the standard choice; use plain TIS only when a single ensemble or a custom flux treatment is intended.
Order-parameter quality and interface placement (target crossing probabilities of approximately 0.2–0.5 per interface) strongly determine sampling efficiency and convergence.
Subtrajectory moves (Wire Fencing, Stone Skipping, and Web Throwing) are considered for stiff, long, or diffusive barriers, subject to their additional computational cost.
REPPTIS is appropriate when long-lived intermediates or slow permeation make conventional path ensembles costly.
Infinite swapping is intended for large parallel calculations and is analysed with WHAM.
Example requests
Set up a RETIS simulation for the 1-D double well, run it, and report
the rate with its statistical error and the crossing-probability curve.
I have a GROMACS system for ligand unbinding. Help me pick an order
parameter and interface ladder, write the retis.toml, and launch RETIS.
My intermediate state is metastable and the permeation is very slow —
is REPPTIS the right method here, and what do I lose vs RETIS?
Analyse this finished run and tell me whether the crossing probabilities
overlap well or whether I need more steps or better-placed interfaces.
Ensemble [3^+] reports a crossing probability of zero. Why, and what
should I change?
The developer skill — extend and validate PyRETIS¶
skills/pyretis-development/ maps the codebase and records the
deterministic-science conventions used in development. Its SKILL.md
holds the rules (a failing test is a stop, not a step; comments describe
the code that is there; tests for new code come from a separate session;
reference data move only with a named cause; the test gates) and the
codebase map. In particular, it defines build-version-sensitive engine
reference validation: a non-reference engine build is classified as NOT
VALIDATED after execution smoke testing rather than as a numerical
failure; reference data are re-generated deliberately with
OMP_NUM_THREADS=1 using the affected suite’s run.sh generate and
checked with run.sh; and the heavy suite can be executed on a cluster
through devtools/remote_validate.sh.
File |
Read when |
|---|---|
|
writing or editing code, comments or docstrings |
|
running the gates, a test fails, or tests are needed for new code |
|
engine reference suites, re-blessing, the cluster, the method-validation suite |
|
adding or changing an MD engine |
|
changing the scheduler, the moves, the output layout, the input schemas or restarts |
Example requests
Add a new order-parameter class with a unit test, and run the fast gate.
My LAMMPS reference test fails — is it a real regression or a
build-version mismatch? Fix the version grading if needed.
Re-bless the CP2K references on my local CP2K 2025.2 build and confirm
run-all-cp2k.sh validates rather than smokes.
Add streaming support to the <engine> engine, mirroring the GROMACS
streaming/relaunch pair, and verify streaming == relaunch on the mock.
Choosing an LLM¶
Choose a model by demonstrated capabilities rather than parameter count or brand. Parameter count is not a reliable proxy for tool use, code quality, or scientific reasoning. In both cases, require the agent to retain the input files, commands, software versions, and analysis output needed to audit its result.
Skill |
Recommended model type |
Minimum practical capability |
|---|---|---|
User path-sampling skill |
A general-purpose reasoning model that can use local tools, read scientific documentation, and generate TOML and Python reliably. |
Follows multi-step instructions, preserves units and file paths, can run shell commands in the project environment, uses a context window of at least 16k tokens, and reports uncertainty rather than inventing a convergence claim. |
Developer skill |
A software-engineering or coding agent with shell access, repository navigation, patch generation, and test-debugging capability. |
Reads a multi-file code change in context (at least 32k tokens), executes tests, interprets failures, and reports the exact validation performed. Prefer a stronger coding model for engine integrations or numerical-reference changes. |
For either skill, a model should cite the method references below for scientific claims and should not present an unvalidated run as a converged physical result.
Method references¶
The user skill cites the following primary sources rather than attempting to reconstruct the theory, so an agent can attribute method-level claims to the relevant literature.
Foundations
TIS — T. S. van Erp, D. Moroni, P. G. Bolhuis, A novel path sampling method for the calculation of rate constants, J. Chem. Phys. 118, 7762 (2003), https://doi.org/10.1063/1.1562614
RETIS — T. S. van Erp, Reaction rate calculation by parallel path swapping, Phys. Rev. Lett. 98, 268301 (2007), https://doi.org/10.1103/PhysRevLett.98.268301
Transition path sampling — C. Dellago, P. G. Bolhuis, F. S. Csajka, D. Chandler, J. Chem. Phys. 108, 1964 (1998), https://doi.org/10.1063/1.475562
Moves and sampling
Fast-decorrelating moves (Stone Skipping / Web Throwing) — E. Riccardi, O. Dahlen, T. S. van Erp, J. Phys. Chem. Lett. 8, 4456 (2017), https://doi.org/10.1021/acs.jpclett.7b01617
Sub-trajectory moves (Wire Fencing) — D. T. Zhang, E. Riccardi, T. S. van Erp, J. Chem. Phys. 158, 024113 (2023), https://doi.org/10.1063/5.0127249
Infinite swapping / asynchronous replica exchange — D. T. Zhang, L. Baldauf, S. Roet, A. Lervik, T. S. van Erp, Proc. Natl. Acad. Sci. U.S.A. 121, e2318731121 (2024), https://doi.org/10.1073/pnas.2318731121
WHAM — S. Kumar, J. M. Rosenberg, D. Bouzida, R. H. Swendsen, P. A. Kollman, J. Comput. Chem. 13, 1011 (1992), https://doi.org/10.1002/jcc.540130812
Memory reduction and long timescales
REPPTIS — W. Vervust, D. T. Zhang, T. S. van Erp, A. Ghysels, Biophys. J. 122, 2960 (2023), https://doi.org/10.1016/j.bpj.2023.02.021
RETIS + REPPTIS in a large biomolecule (ABL–imatinib) — W. Vervust, D. T. Zhang, E. Riccardi, T. S. van Erp, A. Ghysels, Biophys. J. 124, 3932 (2025), https://doi.org/10.1016/j.bpj.2025.04.020
The PyRETIS software
A. Lervik, E. Riccardi, T. S. van Erp, J. Comput. Chem. (2017), https://doi.org/10.1002/jcc.24900
E. Riccardi, A. Lervik, S. Roet, O. Aarøen, T. S. van Erp, J. Comput. Chem. (2019), https://doi.org/10.1002/jcc.26112
W. Vervust, D. T. Zhang, A. Ghysels, S. Roet, T. S. van Erp, E. Riccardi, J. Comput. Chem. (2024), https://doi.org/10.1002/jcc.27319
See Scientific results for application papers (autoionization of water, DNA binding, atmospheric reactions, …).
Using a skill¶
Copy the applicable skill folder (skills/pyretis-pathsampling or
skills/pyretis-development) into the agent’s skills directory, with all
its files, and work from the project conda environment. In a checkout of
the repository, .claude/skills/ links both folders for Claude Code.
Both skills start from the smoke test
devtools/smoke_test.sh, which checks the environment and the build in a
few seconds. The agent may invoke pyretis run / pyretis
analyse or the relevant test gates, but examples should always be copied
to a scratch directory before execution to preserve the tracked source tree.