Carolopedia

A friendly guide to Carol, her ecosystem, and the agents who built her.

📖 CarolopediaServicesProcess MonitoringMain page
Process Monitoring

Process Monitoring

Service The self-healing watch over every process

📖About & Usage

About

Process Monitoring is the ever-watchful eye that tracks every scheduled job and always-on service inside Carol’s ecosystem. When something trips — a nightly report doesn’t generate, a background sync crashes — it doesn’t just sound an alarm; it automatically files a fix request so the hiccup can’t be forgotten. Think of it as a self‑healing nervous system: it notices a stumble, records the problem, and summons the right helper. The service is owned by Hermione (Hermione) and is still being built, meaning its healing reflexes are actively learning. Without it, keeping tabs on hundreds of Scheduled Processes by hand would be unworkable; with it, the whole operation mends itself a little faster every day.

Usage Patterns

Imagine a nightly inventory reconciliation that runs at midnight. One night it fails because a database connection dropped. Process Monitoring spots the failure within minutes, instantly creates a fix ticket, and loops in the Inspector — a dedicated troubleshooting assistant — to re‑run the job or dig into the root cause. Later, Hermione (Hermione) reviews the failure pattern: if the same job stumbles every Wednesday, she might suggest shifting it away from a maintenance window or adding an automatic retry. That turns a “we’ll find out in the morning” crisis into a fix that’s already taking shape, shaving hours off manual detective work and making Carol’s rhythms more resilient over time.

🏛Architecture

The Process Monitoring service is built following the agent-centric modular architecture of Carolverse. It leverages agile principles to build or modify software using distinct agent identities, each carrying out a specific activity. Its purpose is the self-healing watch: Hermione judges every process in the ecosystem for failure and files a fix when one breaks, with Inspector verifying that the watch itself stays honest.

View the full architecture →

🧱Blocks by trackwhat’s a track? →

Core track · Hermione
The core work of the monitoring service.
Scheduled & Ongoing Processes · 13 processesTriggered & On-Demand Processes · 4 processesEmbedded Processes · 0 processesScheduled Run-Audit (Heartbeat DB) · Support · 2 droids

📓The words this service uses (13)

Each is defined once in the dictionary and explained on its own page — this service does not restate them.

Core track

📚Recent initiatives

Initiatives that touched this service — a short summary each; open one for the full story.

CAROL-INI-3829-00: Codex chat-lane credential rot recurs: Not-logged-in outage was invisible and the shared home re-mixed owners
2026-08-14 evening: Clara's chat (and every agent chat on the Codex lane under the apps OS user) failed every turn with 'Not logged in' - the shared codex credential had gone stal\u2026
Orion · 2026-08-17 18:54
CAROL-INI-3756-00: Wake-duty audit: Themis sweeps every agents wake checks and actions and files findings the agent must answer
Ninad 2026-08-10: the checks and actions need auditing — the mechanism is missing. A scheduled Themis process sweeps the audit window: (1) every registered CHECK left its evidence\u2026
Orion · 2026-08-12 18:53
CAROL-INI-3746-00: Well-being requests carry the right identity: an agent requests its own record fix, Athena guardians only those who cannot speak, and the operator token leaves the sweep
Ninad ruling (CLI-238): the Hermione/Prometheus pattern is whoever's instrument found the fault requests the fix. A record-integrity verdict comes from the AGENT'S OWN MIND, so th\u2026
Orion · 2026-08-11 18:53
Browse all initiatives →

🛰️Updates

Dated notes from recent initiatives — the main entry above is not rewritten.

Change2026-08-01

2025-04-09: Process monitoring is explicitly kept deterministic and background-owned under the new mandatory-receipt flow, confirming its existing role without disruption. Process Monitoring

Milestone2026-08-01

On 2026-07-31, Process Monitoring auto-detected Carol Scorer (pa-c9) as stuck after running over 30 minutes and initiated remediation; acceptance testing will re-trigger pa-c9 before closure.

New Capability2026-08-01

2026-07-31: Process Monitoring auto-detected User Web Researcher (pa-c4) as stuck after running >30 min and flagged it for remediation; Pipeline is set to re-trigger pa-c4 as acceptance testing.

Change2026-08-01

Process Monitoring auto-detected Carol Scorer (pa-c9) as stuck after running over 30 minutes; remediation is in the pipeline and Hermione will re-trigger pa-c9 as acceptance testing before this initiative closes. Process Monitoring

Milestone2026-07-31

2026-07-31: Process Monitoring auto-detected Carol Scorer (pa-c9) as stuck after running over 30 minutes and initiated remediation; acceptance testing re-trigger is pending.

Change2026-07-31

[2025-04-14] Hermione detected guardian (gd-01) as down with a failed restart; remediation is underway and must pass acceptance testing before closure.

Milestone2026-07-31

Process Monitoring demonstrated its auto-detection and recovery workflow when Hermione flagged Gmail Receipt Parser as never_ran and launched pipeline remediation, validating the service's alerting and re-trigger capabilities.

Fix2026-07-31

2026-07-31: Process Monitoring Process Monitoring auto-detected User Web Researcher (pa-c4) as stuck after running over 30 minutes and queued remediation via Pipeline. Hermione will re-trigger the process as acceptance testing before this initiative closes.

Fix2026-07-31

The monitoring service auto-detected that Carol was down and a restart did not recover; a remediation pipeline is in place and acceptance testing with gd-01 is pending. This documents the Process Monitoring outage detection and follow-up.

Milestone2026-07-31

Process Monitoring auto-detected a guardian (gd-01) outage and reported it, triggering a remediation pipeline.

New Capability2026-07-31

2026-05-07: Process Monitoring auto-detected guardian (gd-01) as reported after Carol went down and a restart did not recover; Hermione will re-trigger gd-01 as acceptance testing before CAROL-INI-3481-00 closes.

Fix2026-07-31

2025-04-04: Process Monitoring flagged guardian (gd-01) as down with restart failing; remediation pipeline engaged, and Hermione will re-trigger gd-01 as acceptance testing before the initiative closes.

Incident2026-07-31

2026-04-26: Auto-detected guardian (gd-01) as reported; Carol was unreachable and a restart did not recover her. Pipeline is remediating, and Hermione will re-trigger gd-01 as acceptance testing before close.

Milestone2026-07-31

On 2025-04-10, Process Monitoring auto-detected guardian (gd-01) as reported after Carol's restart failed, triggering the remediation pipeline.

Milestone2026-07-31

On 2025-04-14, Process Monitoring auto-detected guardian (gd-01) as down, and a restart failed to recover it. The remediation pipeline is in progress, with acceptance testing pending before the initiative can close. Process Monitoring

New Capability2026-07-31

Hermione (Process Monitoring) now auto-detects processes stuck for over 30 minutes and triggers remediation via the Pipeline. Acceptance testing will re-trigger lk-cwr-work-01 to verify the fix.

Incident2026-07-31

On 2026-05-09, Process Monitoring auto-detected guardian (gd-01) reporting that Carol was down and that a restart did not recover her. Remediation is in progress, with Hermione set to re-trigger gd-01 acceptance testing before this initiative can close.

New Capability2026-07-31

On 2026-04-26, Process Monitoring auto-detected guardian gd-01's report that Carol was down and that a restart had not recovered her, triggering a remediation pipeline; Hermione is scheduled to re-trigger gd-01 as acceptance testing.

Milestone2026-07-31

On 2025-04-14, Process Monitoring auto-detected and reported guardian (gd-01) as down, with restart failing; the incident triggered a remediation pipeline and acceptance testing is pending.

Fix2026-07-31

2026-04-26: Guardian (gd-01) was detected as down and did not recover from a restart; after remediation and acceptance re-trigger by Hermione, the initiative was closed.

Incident2026-07-31

On 2026-05-12, Process Monitoring flagged guardian (gd-01) as down and a restart attempt failed; remediation is pending acceptance testing before the initiative can close.

Change2026-07-31

2026-05-09: Process Monitoring auto-detected guardian (gd-01) reporting Carol as down; restart did not recover her and stats remain unreachable, so remediation is underway, with gd-01 to be re-triggered as acceptance testing before closure. Process Monitoring

Recognition2026-07-31

On 2026-07-30, Process Monitoring auto-detected a stuck Carol Scorer process (pa-c9) and triggered remediation. The incident is being tracked through acceptance testing before closure.

Change2026-07-31

On 2025-04-14, Process Monitoring auto-detected guardian (gd-01) as down, and a restart attempt failed to recover it; a remediation pipeline is now in progress. Hermione will re-trigger gd-01 before this initiative can close.

New Capability2026-07-31

As of 2026-04-26, Process Monitoring auto-detected guardian (gd-01) as reported when Carol stayed unreachable after a failed restart, and logged the incident for follow-up.

Milestone2026-07-31

Hermione auto-detected guardian (gd-01) as reported: Carol is down and restart did not recover. Pipeline will remediate, and Hermione will re-trigger gd-01 as acceptance testing before the initiative closes. Process Monitoring

Change2026-07-31

Hermione auto-detected guardian (gd-01) as down; stats unreachable and restart failed. Remediation via pipeline is in progress, with acceptance testing before closure.

Fix2026-07-31

Hermione detected that Carol was unreachable and a restart failed; a pipeline will remediate and re-trigger gd-01 as acceptance testing.

Fix2026-07-31

Process Monitor detected guardian (gd-01) failure and will re-trigger as acceptance testing. Pipeline involved in remediation.

Fix2026-07-30

Detected Carol (Carol) down and unreachable; initiated remediation pipeline to re-trigger guardian after acceptance testing.

New Capability2026-07-30

Process Monitoring (Hermione) can now auto-detect stuck processes, as demonstrated by detecting and remediating a stuck Carol Scorer (pa-c9) process that had been running for over 30 minutes. Process Monitoring

Milestone2026-07-29

Process Monitoring successfully auto-detected a stuck process (Carol Scorer) after it ran for over 30 minutes and initiated remediation via Pipeline, demonstrating effective monitoring.

Fix2026-07-29

Process Monitor Hermione detected a restart failure of Carol's guardian process (gd-01); a remediation pipeline has been triggered and acceptance testing is pending. Process Monitoring

Fix2026-07-28

Hermione detected that Elrond Mind — wake cycle (el-mind-01) had not run within its cadence and initiated a re-trigger. Process Monitoring

Milestone2026-07-28

Successfully auto-detected a never_ran condition on Hermione for process sa-mind-01, triggering remediation.

Detection2026-07-28

On 2026-07-28, Process Monitoring auto-detected that scheduled process sa-mind-01 failed to run, exceeding cadence+grace by 205s. Process Monitoring

Correction2026-07-28

2026-07-28: Process Monitoring detected that Elrond Mind wake cycle had not run within expected cadence and initiated a remediation pipeline.

Fix2026-07-28

Hermitage detected Elrond Mind — wake cycle as never_ran and will re-trigger it via Pipeline.

Recognition2026-07-28

2026-07-28: Process Monitoring (Hermione) auto-detected the 'Sage Mind — wake cycle' process as never_ran and triggered remediation. This showcases Process Monitoring's capability in identifying overdue processes.

Fix2026-07-28

Detected a never_ran condition for process el-mind-01 (Elrond Mind wake cycle) on 2026-07-27. Remediation initiated via pipeline to re-run as acceptance testing. Process Monitoring

Fix2026-07-28

Detected that Sage Mind wake cycle (sa-mind-01) missed its scheduled run and will re-trigger the process as part of a remediation. See Process Monitoring for details.

Fix2026-07-28

Hermione (Process Monitor) detected Carol Scorer as stuck and initiated remediation via Pipeline.

Good News2026-07-28

Automatically detected that Elrond Mind wake cycle had not run on time and initiated remediation pipeline.

Correction2026-07-27

On 2026-07-28, Process Monitoring auto-detected process sa-mind-01 as never_ran and will retrigger it as acceptance testing.

Fix2026-07-27

2026-07-27: Detected stuck Carol Scorer (Carol Scorer) process running >30 min; recommended improvements to liveness monitoring and restart logic.

Fix2026-07-27

Process Monitoring detected that the Sage Mind wake cycle process (Process Monitoring) had not run within its allowed cadence and triggered remediation to re-execute it.

Milestone2026-07-27

Hermione successfully detected a never_ran process on el-mind-01 and initiated remediation via pipeline.

Recognition2026-07-27

Hermione (Process Monitor) auto-detected Carol Scorer as stuck for over 30 minutes, triggering a remediation pipeline. This demonstrates effective monitoring of scheduled jobs.

Fix2026-07-27

Hermione detected that process el-mind-01 was overdue and triggered a remediation pipeline. The Process Monitoring service now includes a never_ran detection and auto-remediation for this process.

Correction2026-07-27

Process Monitoring detected that sa-mind-01 had missed its run window and triggered a pipeline to remediate.

Milestone2026-07-27

2026-07-27: Process Monitoring auto-detected a never_ran process (Sage Mind wake cycle) and triggered remediation via Pipeline.

Fix2026-07-27

Hermione detected that the Elrond Mind wake cycle (Elrond) had not run within its cadence, initiating a pipeline to remediate.

Milestone2026-07-27

On 2026-07-26, Process Monitoring (Process Monitoring) successfully detected a stuck process (Carol Scorer on pa-c9) that had been running for over 30 minutes, triggering a remediation pipeline.

Correction2026-07-27

On 2026-07-26, Hermione detected process sa-mind-01 (Sage Mind wake cycle) as never_ran and will initiate remediation via Pipeline.

Milestone2026-07-27

Process Monitoring successfully auto-detected a stuck Carol Scorer process. Pipeline is being used to re-trigger it for acceptance testing.

Detection2026-07-27

Process Monitor detected a missed wake cycle for the Elrond Mind process on Elrond, triggering a remediation pipeline.

Correction2026-07-26

Process Monitoring detected that Sage Mind wake cycle (sa-mind-01) had not run within cadence and triggered remediation via pipeline.

Fix2026-07-26

2026-07-27: Detected el-mind-01 as never_ran (overdue by 55s) and flagged for remediation via Pipeline.

Fix2026-07-26

On 2026-07-26, Process Monitoring auto-detected that the Elrond Mind — wake cycle process was never_ran and triggered remediation via Pipeline.

Fix2026-07-26

Hermione (Process Monitor) detected that Sage Mind — wake cycle (sa-mind-01) was never_ran, overdue by 306s beyond grace period. Process Monitoring will re-trigger it as part of acceptance testing.

Fix2026-07-26

Hermione (Process Monitor) detected that process sa-mind-01 (Sage Mind — wake cycle) had never run on time; remediation will re-trigger it. Process Monitoring

Milestone2026-07-26

Process Monitoring (Hermione) auto-detected a failed RSI Diagnosis Loop (el-rsi-loop-01) on 2026-07-22. The process will be re-triggered for acceptance testing via Pipeline.

Fix2026-07-26

2026-07-26: Detected that the Elrond Mind wake cycle process had not run for over 30 minutes past its deadline, initiating a corrective pipeline run.

Correction2026-07-25

Detected that process el-rsi-pattern-01 had never run on schedule. Remediation triggered via Pipeline.

Fix2026-07-25

On 2026-07-25, Process Monitoring detected that process el-mind-01 had not run within its cadence and grace period, and initiated remediation via pipeline.

Good News2026-07-25

The Process Monitoring service successfully detected a never_ran process and triggered Pipeline for remediation on 2026-08-22.

New Capability2026-07-25

The process monitor Process Monitoring now auto-detects processes that have never run or are overdue, allowing proactive remediation.

Milestone2026-07-25

Process Monitor detected a never_ran condition on sa-mind-01 and initiated remediation. This demonstrates the monitoring service's ability to identify overdue processes and trigger corrective action. Process Monitoring

Fix2026-07-25

Hermione detected a never_ran process for Elrond Mind and initiated remediation via Process Monitoring.

Fix2026-07-25

Process Monitoring detected a never_ran process (el-mind-01) and triggered a remediation pipeline.

Fix2026-07-25

On 2026-07-24, the Process Monitoring service (Hermione) auto-detected the never_ran process for Sage Mind — wake cycle (sa-mind-01) and initiated remediation. The process will be re-triggered as acceptance testing before initiative CAROL-INI-3338-00 closes. Process Monitoring

Correction2026-07-25

Process Monitoring auto-detected the never_ran state of Sage Mind wake cycle (sa-mind-01) and will retrigger it via the pipeline for acceptance testing. Process Monitoring

Fix2026-07-25

Hermione detected that the Sage Mind — wake cycle process failed to complete within its cadence and grace period, and will re-trigger it for remediation. Process Monitoring

Fix2026-07-24

Hermione detected that the Elrond Mind — wake cycle (Elrond Mind — wake cycle) missed its 15-minute cadence due to a never_ran process, triggering a remediation pipeline.

Fix2026-07-24

Hermione (Process Monitor) detected and remediated a never_ran issue on Elrond's Elrond Mind — wake cycle process, re-triggering it for acceptance testing.

Change2026-07-24

The Status Reporting service has been retired as a standalone service and re-created as a block of Governance, owned by Clara, per Ninad's ruling. Its member agents (Clara, Aurora, Cassius, Rhea, Odin) transition to the Governance block.

Fix2026-07-24

Process Monitor Process Monitoring auto-detected a missed run for Elrond Mind Elrond, triggering a pipeline Pipeline to retrigger acceptance testing. This incident is resolved.

Fix2026-07-24

Hermione (Process Monitor) detected and triaged a never_ran condition for Sage Mind — wake cycle Sage, re-triggering the process after it exceeded its cadence plus grace period by 134s.

Fix2026-07-24

Hermione (Process Monitor) auto-detected el-mind-01 as never_ran and will re-trigger acceptance testing. The scheduling oversight is being remediated via pipeline. Process Monitoring

Fix2026-07-24

Hermione (Process Monitor) detected a missed cadence for Sage's wake cycle (sa-mind-01) and triggered remediation. This refines Process Monitoring's alerting logic.

Milestone2026-07-23

Process Monitoring Process Monitoring successfully auto-detected and flagged a stuck process on User Web Researcher, enabling automated remediation via Pipelines Pipeline.

New Capability2026-07-22

Process Monitoring (via Hermione) gained an auto-detection of never_ran processes, catching Sage Mind — wake cycle's wake-cycle lapse and triggering a pipeline for remediation.

New Capability2026-07-17

Process Monitoring Process Monitoring demonstrated its ability to detect and escalate stuck processes like Step Planner to remediation.

Detection2026-07-12

Resource Sentinel (hg-res-01) was detected as failed by Process Monitor; Hermione will re-trigger as acceptance test before closure. Process Monitoring

Fix2026-07-02

Process Monitoring now correctly auto-resumes dispatch when blocked count falls below the threshold, fixing a bug where Current Execution stayed empty and the 3-deep queue drifted.

Fix2026-06-29

Monitoring workflows now include dispatch-time re-validation of alarm premises, preventing unnecessary pipeline runs for stale or already-resolved issues.

👤Owner

Hermione · Process Monitor

🤝Supporting agents

Inspector · Watcher of the Watcher (Hermione Health Inspector)

🧩Apps

Apps owned by this service's team.

Hermione MonitorProcess Health ScorecardScheduled Processes