{"wiki":{"id":1385,"slug":"he-llmhealth-01","entity_type":"droid","entity_id":"he-llmhealth-01","title":"LLM Provider Health Watch","prose_md":"## About\n\nLLM Provider Health Watch is a droid that quietly checks whether every LLM provider lane is still working — so your conversations don't suddenly fail with an `insufficient_quota` error because an account ran dry or a service went down. It reads health‑status latches every 15 minutes, and once a day it makes one tiny probe call per API‑key provider (like a single‑token request) to verify the lane is really alive. If a lane is dead, it sends exactly one critical alert via the shared alert bus and marks the outage flagged, so the right people know immediately.\n\nThe droid never probes the Claude CLI lane (that lane is fleet‑disconnected by a standing policy), and it resets the alert latch automatically the first time the provider answers successfully again. It always reports a summary of what it found in the run‑audit summary, which feeds into tools like [[audit-scorecard]] and [[monitoring]]. All its checks run on a [[scheduled-processes]] cadence, so no human has to remember to verify account balances.\n\nLLM Provider Health Watch is owned by [[agt_016]] and is not a one‑off — it’s one of 209 identical droids in the unnamed `(none)` family that many agents run, including [[agt_008]], [[agt_016]], [[agt_030]], and others. Because it calls an LLM for that daily probe, it is not pure software: it relies on Claude (model unspecified) to perform the tiny test call, which is the very thing it’s guarding.","namesake_json":"{\"engine\": \"claude\", \"model\": null, \"claude_purpose\": \"tiny probe call to confirm LLM provider is alive\", \"family\": \"(none)\", \"family_size\": 209, \"is_unique\": false}","profile_pic_path":"","source_hash":"fa92e5ecb4434141f5088d2d685ea048f34b73815e52851d05f52a0ef0257d9e","status":"active","last_generated_at":"2026-07-28 03:33:35","created_at":"2026-07-28 03:33:20","updated_at":"2026-07-28 03:33:35"},"facts":{"id":"he-llmhealth-01","name":"LLM Provider Health Watch","machine_name":"he_llmhealth_01","owner":"agt_016","function":"Watches every LLM provider lane and alerts once when one dies","process_type":"scheduled","schedule":"Every 15 min","process_name":"llm-provider-health-watch","avatar_color":"#94a3b8","created_for":"CAROL-INI-2863","purpose":"Reads the per-provider balance-health latch every 15 minutes and makes one tiny probe call per provider per day, so an exhausted or broken account (like the 2026-07-15 OpenAI insufficient_quota outage) is discovered by monitoring, not by Ninad mid-conversation.","duties":"Check shared/llm_health provider latches; daily 1-token probe per API-key provider; on a down lane send ONE critical alert via shared/alerts and mark it alerted; report lane status in the run-audit summary.","constraints":"Never probes the claude CLI lane (fleet-disconnected by CAROL-INI-2389); alerts once per outage, latch clears on first successful call.","status":"running","gender":"","archetype":"","building_block":"mon_scheduled","service_override":null,"enabled":1,"task_key":"monitoring.watch_pass","model_free":0},"page":{"type":"droid","page_class":"main","class_label":"Main page","kind_label":"Droid","kind_gloss":"","listed":true}}