Continuously Improving Agents with Langfuse — Annabell Schäfer & Lotte Verheyden, ClickHouse
A frustrated user types in all caps and tells an IT support agent that its answer did not work. Those interactions become signals that a team can find in production traces, rather than anecdotes lost in a chat history. Lotte Verheyden and Annabell Schäfer build that monitoring workflow with Langfuse and a sample support agent called Specs. The application helps a dad with phone questions, and the workshop follows a request through agent observations, model calls and tools to show what a trace needs to contain before it can support useful evaluation. The hands on setup connects a TypeScript application built on the OpenAI Agents SDK to Langfuse. Schäfer adds a judge evaluator for user disagreement, a code evaluator for all caps and a scope check that compares the user's request with the system prompt. Each evaluator targets the relevant messages through explicit variable mappings, so its score answers a defined question. Seeded production traces then give a coding agent enough evidence to investigate patterns through the Langfuse CLI and skill. The analysis surfaces issues such as unnecessary tool calls, retries, latency and incomplete evaluation coverage, and the presenters explore a proposed code change. The closing discussion places this online monitoring inside a continuous improvement loop: turn observed failures into representative datasets, run offline experiments, check regressions and repeat as real usage changes. Speaker info: Annabell Schäfer: - https://www.linkedin.com/in/annabell-schaefer/ - https://github.com/annabellscha Lotte Verheyden: - https://www.linkedin.com/in/lotteverheyden/ - https://github.com/Lotte-Verheyden Workshop resource: - https://langfuse.com/workshop Timestamps: 0:00 - The AI engineering loop 3:08 - Online and offline improvement 5:12 - Build useful traces 6:38 - Monitoring and user feedback 8:01 - Meet Specs, the dad IT support agent 10:06 - Repository and account setup 12:28 - Run the TypeScript application 13:23 - Inspect a live support trace 15:02 - Choose application specific signals 16:21 - Configure the evaluator model 18:16 - Set up a disagreement judge 19:16 - Map trace messages into the evaluator 21:07 - Code based all caps detection 24:16 - Test and execute the monitors 24:58 - Generate frustration and disagreement 26:13 - Inspect evaluator scores 27:47 - Detect out of scope requests 31:51 - Seed production traces 33:50 - Install the CLI and skill 35:24 - Investigate traces with a coding agent 37:24 - Retrieval, retries, and coverage findings 39:05 - Propose a code change 40:03 - Maintain datasets and close the loop 42:25 - Workshop resources and next steps
More like this

AI Security Engineer Foundations + Certificate — Javier Garza, Snyk

Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI

SonarQube + OpenAI: Agentic Development — Killian Carlsen-Phelan, Sonar

Let Your Agent Cook: Using Skills to Evaluate and Improve Your App — Ankur Duggal, Arize AI
Join the discussion
Sign in to join the discussion
Sign in