๐What it means
How well an agent performs one of its competencies, judged every day from its real attempts and never from a practice test: fluent, proficient, learning, forgotten or ungraded. Athena's Competency Grader records one grade per agent per competency per day with the evidence it rests on; rows are never changed, so the history is kept. Refusals by a switch, a budget or a gate are not counted, and an outcome that cannot be read is recorded as unauditable and counted neither way. Where several droids produce one response together, the turn is judged as a whole and the droids inside it share its grade.
๐Easily confused with
[{"term": "Agent score", "distinction": "The agent score rewards CONDUCT against seven criteria; a competency grade says how reliably the agent performs one specific thing it can do."}, {"term": "Weight", "distinction": "Weight is the tokens an agent carries per step; a grade is about outcomes of attempts, not cost."}]
๐Nuance
The window and thresholds are a dated record in the grade store, amendable without a code change.
๐Also called
- grade
- competency grading
๐Filed under
Category: governance.
Service: Agent and Human Resources.
Source: Data Dictionary ยท Public information reflected here.