Teaching Agents to Search with NVIDIA Data Designer — Dhruv Nathawani, NVIDIA
50,000 seed examples become 7,000 training records after generation and filtering. Dhruv Nathawani uses that funnel to explain how NVIDIA builds synthetic data that teaches a model to search, rather than answer from memory. The starting material is a set of walks through Wikidata: connected entities become riddles whose answers require multiple searches. Data Designer turns those seeds into questions, tool calls and responses, with Tavily providing search through MCP. The goal is a reusable pipeline for teaching a chosen behavior, with search serving as the worked example. Three notebooks build the workflow in stages. The first combines seed topics, difficulty samplers, prompt templates, structured outputs and model judges to generate question and answer pairs. The second adds tool access and captures search traces. The third creates search riddles from graph paths and produces complete trajectories for supervised finetuning. Previewing a few records before running a batch makes the process inspectable and easier to revise. Nathawani then examines the difficult part: verifying that generated answers are correct when even the source graph can become stale. Model judgments, deterministic checks and human review contribute different evidence; no single score establishes dataset quality. Questions distinguish seed generation from agent execution, explore reinforcement learning alternatives and discuss other training tasks. The closing examples use synthetic personas to preserve demographic distributions, while the broader lesson remains practical: design for diversity, inspect examples, overproduce, filter and version the recipe. Speaker info: - https://www.linkedin.com/in/dhruvnathawani/ - https://github.com/NVIDIA-NeMo/DataDesigner Timestamps: 0:00 - Turn an agent behavior into a data pipeline 3:10 - Why internet data is not enough 5:12 - Diversity beyond repeated prompting 6:23 - Seed, sampler, and generation columns 8:55 - Notebook one: question and answer data 10:32 - Difficulty samplers and distributions 11:35 - Prompt templates and structured outputs 12:44 - Preview a few records 13:24 - Judge question quality 14:19 - Generate a batch 15:26 - Notebook two: search through MCP 18:00 - Inspect a search trajectory 19:20 - Why models need search 21:39 - BrowseComp and difficult questions 23:11 - Wikidata paths as search riddle seeds 26:40 - Notebook three: end to end search data 28:36 - Tool instructions and example calls 30:32 - Inspect the generated riddles and traces 33:47 - Seed generation versus agent trajectories 34:29 - Supervised finetuning and distillation 37:04 - Reinforcement learning alternatives 39:24 - Filter and evaluate synthetic data 41:04 - Stale sources and answer verification 43:08 - Human review of model judgments 45:39 - Data generation practices and other tasks 47:17 - Synthetic personas and demographic distributions 51:35 - Curation, anonymization, and related tools 53:39 - Resources and persona generation
More like this

AI Security Engineer Foundations + Certificate — Javier Garza, Snyk

Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI

SonarQube + OpenAI: Agentic Development — Killian Carlsen-Phelan, Sonar

Let Your Agent Cook: Using Skills to Evaluate and Improve Your App — Ankur Duggal, Arize AI
Join the discussion
Sign in to join the discussion
Sign in