Create test data
Turn approved knowledge into questions and expected answers.
Dataset guideLLaDAR creates test questions from your knowledge, sends them to an existing Agent, checks the answers, and writes a report. Each stage can use a custom local Skill.
Trying one question in a chat window is hard to repeat or review. LLaDAR connects the source material, test question, Agent response, evaluation result, and report, so you can see what an existing Agent actually did.
Turn approved knowledge into questions and expected answers.
Dataset guideFind a public way to use the Agent, then save its responses.
Run-agent guideCompare each response with what was expected, then calculate the results.
Evaluation guideTurn the saved results into a readable report with limitations and details.
Reporting guideAfter the provider and target Agent are configured, LLaDAR can run the whole workflow automatically. A local directory containing SKILL.md can tailor how questions are generated, answers are judged, or results are summarized.
Run the included support demo Agent from start to finish.
Get startedUse the run-agent guide for a project or controlled test service you already have.
Test an existing AgentUse local Skills to choose records, judge answers, or format a report.
Learn about SkillsThe workflow saves the generated dataset, responses, execution records, evaluation result, and Markdown report. A successful replay shows that LLaDAR reached the selected Agent entry point; it does not by itself prove that every answer is correct.