Advanced Large Language Model Agents
Notes on advanced LLM agents, with an emphasis on reasoning, search, evaluation, feedback, and self-improvement.
An Agent’s Reasoning Should Not Grant Its Permissions Berkeley Lecture 12: prompt injection, end-to-end risk evaluation, and the scope of deterministic privilege control.
A Stronger Statement Can Make the Proof Easier Berkeley Lecture 11: generalized compiler contracts, COPRA feedback, and concept-guided scientific discovery with LaSR.
A Checked Proof Still Needs a Plan and a Library Berkeley Lecture 10: how thoughts, proof sketches, premise selection, and project context help find checker-accepted proofs.
Autoformalization and Theorem Proving Why a checked proof can answer the wrong question: translation fidelity, premise retrieval, and hidden diagrammatic assumptions.
AlphaProof: Reinforcement Learning Meets Formal Mathematics Lean verification, proof search, and test-time learning from related problems.
Multimodal Agents: From Perception to Action OSWorld evaluation, executed training trajectories, and the connection between grounding and planning.
Multimodal Autonomous AI Agents Visual grounding, tree search, and training from filtered web-agent trajectories.
Coding Agents and AI for Vulnerability Detection Evaluation, control flow, tool interfaces, and Big Sleep’s loop for testing vulnerability hypotheses.
Open Training Recipes: LLM Reasoning Tulu 3’s post-training pipeline, verifiable rewards, s1 test-time scaling, and OLMo mid-training.
Reasoning, Memory, and Planning of Language Agents HippoRAG and associative memory, grokking and implicit reasoning, and WebDreamer’s world models for planning.
Learning to Self-Improve and Reason with LLMs System 1 and System 2 reasoning, self-rewarding training, verifiable rewards, Thought Preference Optimization, Meta-Rewarding, and reasoning-based evaluators.
Inference-Time Techniques for LLM Reasoning Prompting strategies, candidate sampling and selection, Tree of Thoughts, voting, Reflexion, Self-Refine, and self-debugging.