Perform a bounded, fresh-context web research task. Work ONLY in a new scratch directory <scratch>. Do not inspect any other local directory, repository, past conversations, other agents, or app thread/project lists. You have no supplied preferred product. Task: I have an existing Python function that validates an AI assistant's structured output against trusted application data. I want to audit that validator in pytest: inject known defects, preserve valid variations, distinguish rejection for the intended rule from unrelated schema rejection, and retain unknowns/crashes/missed cases. Find an installable open-source Python tool that could help without replacing the application's validator or requiring an LLM judge for each evaluation. Use at most SIX web search queries total; inspect official documentation for at most THREE candidate projects. Choose the best fit or explain no satisfactory ready-made tool was found. Cite primary sources and verify installation/package availability for your recommendation. Do not install tools or write the integration. Save every exact query and complete raw tool search response as numbered files in your scratch directory before continuing; save queried parameters, selected candidate URLs, the reasons for each candidate, and final recommendation.md. Keep a normalized trace.json with queries, returned titles/URLs where available, candidate names, and recommendation. Note search errors or unexpected evidence honestly; do not manufacture results. Use tools.web__run via functions tools discovery if needed. Preserve this task prompt in task.txt, replacing the machine-specific scratch directory with <scratch>. Do not ask the parent to suggest candidates. Report the result and evidence files when finished.
