ServiceNow’s AutoSynthData work describes a way to generate training tasks around agent failures. The useful lesson is that a generated task needs a trustworthy success test, not just a plausible prompt.
The proposed workflow
ServiceNow-AI published its AutoSynthData article on October 2. The approach uses tasks built from a system context, a user request, and a verifier. It describes generating tasks that challenge a target model while a stronger teacher can complete them, then expanding validated examples. The article reports supervised fine-tuning experiments; it identifies reinforcement learning as future work. These are the authors’ results, not an independent replication. ServiceNow-AI article.
The related EnterpriseOps-Gym dataset provides context for enterprise-agent evaluation. It should not be confused with independent proof that AutoSynthData improves every agent.
Test the verifier before trusting the sample.
A verifier can accept an incorrect outcome if it checks only a superficial signal. For example, finding a ticket identifier in an answer does not prove that the right ticket was updated. The test must inspect the intended state and reject a convincing explanation of work that never happened.
Build negative cases deliberately. Give the verifier the wrong record, an incomplete update, and a successful action on an unrelated object. Each should fail. A test that rejects nothing is not a useful training boundary.
A dataset review checklist
- Define the exact final state required for success.
- Check that the teacher’s execution genuinely reaches that state.
- Confirm that incorrect or partial states fail verification.
- Remove private identifiers and sensitive records.
- Track which generated tasks are variants of the same seed.
- Keep evaluation examples separate from training examples.
- Freeze the holdout set before comparing models.
- Review diversity by task type, not just by sample count.
Our Jev guide offers a related discussion of constrained decisions. The AI-content audit guide explains why generated material still requires evidence and review.
Editorial takeaway: more synthetic tasks are useful only when their success conditions are reliable. Validate the test before scaling the generator.