Google Research has introduced ToolGrad, a method for producing reliable datasets that teach models to use tools. Instead of writing a question first and hoping an API chain can solve it, the system builds a valid execution before generating the corresponding user request.

Quick answer

ComponentPurpose
Tool chain firstEnsures a solution path actually exists.
Textual gradientExplains failure and guides correction.
ToolGrad-500A 500-task evaluation set announced by Google.
Validation rate99.8% for generated data in the reported experiments.

The synthetic-data problem

An agent must choose a tool, form parameters, interpret the result and sometimes call another tool. A plausible question does not guarantee that the available APIs can answer it. Impossible or ambiguous examples can teach the model the wrong behavior.

ToolGrad reverses the process. It selects and executes a sequence of tools, obtains an answer, then formulates a user request consistent with that path. Failures become textual critiques that revise the example.

A four-module loop

The system uses a proposer, executors, a selector and an updater. It can attempt several solutions, retain the best ones and use natural-language feedback to guide the next generation. “Gradient” is an analogy here: the API is not mathematically differentiated; a textual diagnosis supplies the direction of improvement.

Google reports lower cost than generate-and-filter approaches and gains when Gemma 3 models are trained on the resulting data. Those results remain specific to the evaluated tools and protocols.

Lessons for engineering teams

Even without reproducing ToolGrad, teams can execute every training example, retain call traces and classify failures: wrong tool, invalid parameter, missing data or unverifiable answer. That discipline improves evaluation as much as training.

Synthetic data also needs checks for leaked secrets, dangerous operations and linguistic diversity. A high task success rate does not automatically measure safety.

A central agent bottleneck

Models are advancing quickly, but high-quality tool-use data remains expensive to create and verify. ToolGrad suggests that starting from an executable trajectory is more useful than collecting polished but impossible conversations. For production agents, dataset quality now depends as much on the test harness as on the selected model.