{"wiki":{"id":2507,"slug":"pe-dev-01","entity_type":"droid","entity_id":"pe-dev-01","title":"Post-Execution Evaluator","prose_md":"## About\n\nPost-Execution Evaluator checks finished work to confirm it actually succeeded. Owned by [[agt_003]], it verifies the original success criteria and test plan, looks for missing or broken functionality, and produces a detailed pass/fail report.\n\nIt checks frontend code, API paths, and critical behavior in a browser. It reports problems rather than fixing them and escalates critical or repeated failures. Its exact name is unique, but it belongs to a standard reviewer family with 31 droids used across several agents. It is not pure software: it calls Claude’s opus model to evaluate completed tasks against their predefined criteria and tests.","namesake_json":"{\"engine\": \"claude\", \"model\": \"opus\", \"claude_purpose\": \"evaluate completed tasks against success criteria and test plans\", \"family\": \"reviewer\", \"family_size\": 31, \"is_unique\": true}","profile_pic_path":"","source_hash":"5af30ed73ed6e5492e7f7cfeaf5826dc5c24f82ffb4e5838188fab343076241d","status":"active","last_generated_at":"2026-08-04 00:43:47","created_at":"2026-08-01 03:27:33","updated_at":"2026-08-04 00:43:47"},"facts":{"id":"pe-dev-01","name":"Post-Execution Evaluator","machine_name":"PE-DEV-01","owner":"agt_003","function":"Evaluates completed task against success criteria and test plan","process_type":"triggered","schedule":"On demand","process_name":"post_exec_evaluator","avatar_color":"#f97316","created_for":"Argus is the tester — their job is to catch bugs and gaps in completed work. This droid helps by checking whether finished tasks actually met their success criteria and passed their test plans, giving Argus a clear verdict before the work ships.","purpose":"When a task finishes, this droid verifies it actually worked. It checks the success criteria, validates the tests, and confirms the frontend still functions. Argus gets a detailed pass/fail report with specifics on any failures.","duties":"- Verifies each success criterion in the completed work\n- Checks whether the test plan passed\n- Looks for any broken or missing functionality\n- Tests the frontend in two ways: code-level checks (syntax, API paths) and in-browser checks (app loads, pages display correctly)\n- Creates a report listing what passed, what failed, and why\n- If critical problems arise that the droid can't handle, notifies Argus's manager and specialist agents","constraints":"- Only checks against criteria that were set before work started; does not define new tests\n- Identifies problems but does not fix them\n- Frontend tests focus on critical issues only (broken code, bad paths, app unresponsive, missing content)\n- Stops trying and escalates if the evaluation service fails repeatedly\n- Only checks technical correctness and task completion; does not assess design, user experience, or documentation","status":"running","gender":"male","archetype":"reviewer","building_block":"review_step","service_override":null,"enabled":1,"task_key":"initiatives.tests_verified","model_free":0},"page":{"type":"droid","page_class":"main","class_label":"Main page","kind_label":"Droid","kind_gloss":"","listed":true}}