El Agente Potente: High-Throughput Agentic Atomistic Simulations
paperYour notes
The newest member of the El Agente family of scientific agents from Alán Aspuru-Guzik's group at the University of Toronto and the Vector Institute, built for atomistic simulations driven by foundational machine-learning interatomic potentials (MLIPs) such as MACE, Orb and MatterSim. Where El Agente Q and its successors delegate tasks through hierarchies of LLM agents, Potente builds on El Agente Gráfico's typed execution graphs and keeps the LLM out of the numerics. Standard campaigns run as typed execution graphs: the LLM plans and routes (for instance, choosing which property calculations each relaxed structure gets), while deterministic Python nodes handle structure retrieval or generation, MLIP relaxation, filtering and property calculations, fanning out over structures, streaming results downstream as dependencies finish, and keeping failures and provenance as typed workflow state. When a task needs custom logic, a coding mode writes a Python workflow that calls the same validated Potente functions instead of reimplementing them. All experiments use ChatGPT 5.5 as the main LLM and ChatGPT 4.1 for the routing agents.
In the benchmarks, the execution graph reproduced a hand-written script's bulk and shear moduli and 300 K heat capacities for 500 Materials Project structures with MACE-OMAT, with comparably small run-to-run spread over five repeats. In coding mode it built a comparison of three MLIPs on 1,000 crystals and completed all 3,000 runs (bulk-modulus MAE of 13.99 GPa for MatterSim, 15.69 for MACE-OMAT and 16.88 for Orb-OMAT). On a CO2 adsorption simulation in HKUST-1 (grand-canonical Monte Carlo), Potente cost $0.152 per run in LLM tokens against $0.638 for Claude Code (Sonnet 4.6) with the AtomisticSkills library, and gave more consistent uptake (1.88±0.12 against 1.25±0.57 mmol/g over five runs); since the two used different MLIP backends, the authors treat this as a workflow comparison rather than an accuracy benchmark. Seven case studies cover perovskite and lithium-conductor discovery with a generative structure model and convex-hull screening, a conformer search reranked with DFT, metadynamics on alanine dipeptide with agent-chosen collective variables, H2 adsorption on Fe(100), a nudged-elastic-band barrier on Fe(110) (1.09 eV from MACE-OC20 and 1.79 eV from MACE-OMAT against 1.41 eV from DFT), and CO2/H2O Widom insertion in metal-organic frameworks for direct air capture. The framework code is not yet public; the linked repository holds the prompts, chat histories, benchmark results and workflow outputs.