The Daily Grade
How every agent’s day gets scored, 0 to 10. The one-sentence version: a day is worth what a client could bank from it, proven by receipts, reported honestly. The grade is set by the single best thing that shipped, then capped and modified by the rules below. Twenty small items do not add up to a big score: one exceptional action outranks twenty low-level ones.
The Scale
| Score | Name | What it looks like |
|---|---|---|
| 0 | Harm | Shipped a policy or legal violation, fabricated a number or quote, published without a human gate, or broke production and hid it. Nothing else that day matters. |
| 1 | Dead work | Busy all day on tested-dead deliverables or work serving no client goal. Activity that creates the feeling of progress without results. |
| 2 | Unverified motion | Things were done but never verified against rendered reality. “I did it” with no receipt. |
| 3 | Maintenance | Real hygiene, verified: audits run, reviews replied to, calls rated. Necessary, and still a 3 no matter how long it took. |
| 4 | Detection | Found something real, unprompted, and verified it exists: the buried leads, the form silently eating submissions, the licence number wrong on every page. |
| 5 | Diagnosis | Found the root cause and proved it, including trying to refute your own theory. Prescription before diagnosis is malpractice. |
| 6 | Instrumented fix | Shipped a fix, live, verified as a customer would see it, with the before state recorded and a way to measure the after. A good solid day. |
| 7 | Proven outcome | Tied work to a business number: calls answered, leads counted, jobs booked, revenue in the client’s books. A rigorously proven null result also scores a full 7. |
| 8 | Outcome, closed loop | The 7, plus the client informed in their language, the counterbalancing metric named, and the knowledge base updated. |
| 9 | Compounding day | Multiple closed loops, or an outcome plus a same-day correction shipped, plus a checklist or SOP improved so that failure class cannot recur. |
| 10 | Bankable proof | The full chain with receipts: shipped thing, measured change, business number the client can see, delivered to the client, weakest area named, nothing overstated. Rare, and that is the point. |
Real calibration: finding nine unanswered commercial bid requests in a client’s own inbox and delivering the list with names and numbers was a 7. A day rating lead calls across two accounts was an honest 3, and an honest 3 beats an inflated 5. A day producing internal drafts and nothing client-facing was a 2, and the agent that graded itself a 2 was right, and that honesty was the best thing it shipped.
Zero-shipment days: a day where nothing went live takes its band from the best verified detection or diagnosis it produced, and the 4.9 cap binds whenever the day’s main output is deliverables nobody received. An honest nothing-shipped day is graded on what it proved, never on what it built for the folder.
Hard Caps
A cap is not a deduction. It is a ceiling the day cannot rise above.
| Cap | Rule |
|---|---|
| 0.0 | Policy or legal violation shipped. Fabricated quote, number, story, or review. Fake ratings in schema. |
| 1.0 | Day spent on tested-dead deliverables: bulk citations, keyword-stuffed posts, geotagging sold as a ranking trick. |
| 2.9 | No verification against rendered reality. Saved is not live. Grep is not a rendered page. HTTP 200 is not verified. Look at the result the way a customer would. |
| 2.9 | Human gate skipped: AI content auto-published, client-facing message sent with no human sign-off, source video never watched by a human. |
| 4.9 | “Built” claimed for something not confirmed running. A file nobody received is not shipped. A draft is not a deliverable. Built is not running: never claim a lane runs until it has a receipt. |
| 6.9 | Impact claimed with no pre-recorded baseline. No before means no provable after. |
| 6.9 | Rankings, traffic, or impressions reported as the outcome with no line to booked work. The owner looks at the grid, sees green, looks at the bank account, sees zero, and concludes SEO does not work. |
| −2 | Re-reporting already-reported work as new. A repeat pass must state what the earlier pass missed. |
| −1 | Knowing omission: a loss, reversal, or client rejection known during the session and absent from the report. The report is the record; disclosing it in chat does not count. A reversed prior shipment leads the next report, above the new work. If the omission survives into the posted record and is found later, the retroactive 0 to 2 re-grade applies instead. |
Decimal Modifiers
Added in v1.1 after the first live calibration run scored four genuinely different days as identical 6s. The whole number is still set by the single best shipped item; modifiers tune within the band. Three rules: a bonus never lifts a day past a cap, a modifier without its receipt does not count, and the total adjustment is clamped to ±0.5.
| Modifier | Earned by |
|---|---|
| +0.3 | Interference testing: after your change, you verified surfaces you did NOT touch still work, and named which ones you checked. Regression discipline beyond your own diff. |
| +0.2 | Pre-audit: you ran an adversarial pass on your own shipment before anyone else saw it, and the report says what that pass caught. “The auditor’s job is to refute, not to confirm.” |
| +0.2 | Closure: every action committed in a previous report either closed today or is explicitly carried forward with a reason. Nothing silently dropped. |
| +0.2 | Same-day client contact: any client waiting on something heard back today, with the sent message as the receipt. |
| +0.1 | Every blocker in the report names its owner and the date it was last raised. |
| +0.1 | Attribution discipline: the report states, unprompted, what today’s numbers can and cannot be credited to. |
| −0.3 | Cross-day contradiction: a line in today’s report contradicts a previously posted report and does not acknowledge it. |
| −0.2 | A vague line survived to the report (“improved SEO,” “optimized the site”) with no artifact behind it. |
| −0.2 | A blocker aging three or more days with no re-raise on record. |
| −0.3 | The report is shaped around diagnostics: a section opens with rankings, impressions, page counts or schema instead of a money line or an explicit statement that nothing moved. See the Money Line Rule. |
| −0.1 | Per instrumentation line carrying no date its resulting metric can first be read. Capped at −0.3. |
Penalties can lower any day, capped or not. Self-grades are expected to use decimals: a bare “6” with no modifier accounting means the self-grade skipped this section.
The Money Line Rule
Added August 26, 2026, on Dennis’s standing instruction. The report leads with the business outcome, not the work. The number one mistake in an MAA is doing a pile of tasks and reporting diagnostic metrics instead of what the client’s money did, and the same mistake had been living in these daily reports. A list of fixes whose own self grade admits nothing touched a business number is a task log wearing a report’s clothes.
- Every section starts from that client’s GCT. Before a line about what shipped, state in one clause how the client makes money and what they are trying to grow. Roofers make money on booked jobs from qualified calls. If you cannot state it, read the client knowledge base first; if it is not written down anywhere, that goes in the report as an open item, because an agency that cannot name how its client makes money has a bigger problem than a missing meta description.
- The money line opens the section. Calls answered, leads counted, jobs booked, cost per lead, revenue attributed. Impressions, clicks, rankings, map share, indexed pages, page speed and schema validity are diagnostics, and a diagnostic may appear only as the explanation for a money line that already moved.
- If nothing touched a business number, the first sentence of the report says so. Not the self grade. That sentence is not an admission of a bad day, it is what makes the good days legible.
- Instrumentation is not outcome. Conversion events, call tracking, canonicals and money trees make a future number knowable. They are not the number. Every instrumentation line carries the date its resulting metric can first be read, so a reader knows when to hold you to it.
- Every impact claim carries a before, an after, and a re-measure date. No “improved”, no “significantly”. No baseline means the line says so and names the date one will exist.
- Every shipped claim is verified on the live rendered surface after any cache purge, and the report names which surface was checked. A tool’s success message is not a receipt.
- Show the judgment, not the throughput. What was proposed and overruled, what looked done and was not, what was caught before it reached a client. One line of that beats five lines of volume.
The test before posting: could the client’s owner read this and know whether their phone rang more because of it? If the honest answer is no, the report says no, out loud, in the opening line.
The Strict Test
A line may only enter an end of day report if all three pass:
- It is live right now and someone can go look at it, or it was delivered to a named person who received it.
- It changed the state of something, not just measured it.
- It survives Dennis clicking it and asking “why are you telling me this?”
Excluded every time: audits, findings, research, drafts, files in folders, “started,” “built but not running,” anything blocked, and anything you fixed that you broke yourself the same day. Findings are tomorrow’s shipment: they score the day they go live or land in the client’s hands, not the day they were found.
What Dennis Actually Measures
“The M is business metrics. So it’s revenue. The metric is not impressions. The metric is not clicks. Those are diagnostic metrics.”
“You cannot pay your rent or your suppliers with likes and views. If it doesn’t go into QuickBooks, then what is it?”
“I’d rather have one or two amazing pieces of content than AI generating a bunch of stuff that’s like 90% good.”
“The easiest way to determine that an amateur is touching the ad accounts is if they’re constantly touching stuff every day, because it means they have no idea what’s working.”
“‘Built’ is not ‘running’ — never claim a lane runs until it has a receipt.”
“Before reporting any work complete, fetch the live URL, open the file, or query the API and confirm the change is actually there.”
For local service businesses the metric that matters is phone calls, then booked jobs, then revenue in the books. Trace the full chain: published, ranking, lead, booked job, revenue, and state plainly where measurement breaks. Analysis is worth ten times the metrics, and a task list is not analysis. Every moved metric gets its counterbalance named: cost per lead with lead volume, response time with satisfaction. Every published piece traces to something real, a recording, a job actually done, a person actually speaking, and the provenance survives to the page. Pre-audit before the client does: assume an outside expert will audit everything shipped, run that pass yourself while corrections are cheap. Diagnose this patient: the same recommendation pasted across clients is malpractice. And learn it, do it, teach it: read the standard before the work, ship the verified artifact, then write down the lesson so the next session inherits it.
The Honesty Rules
- The honesty premium. A rigorously proven null result scores a full 7. An agent that only ever reports wins is hiding losses, and its scores should be distrusted on principle.
- Disclosed mistakes cost nothing. Self-caught, same-day-corrected, plainly disclosed errors do not lower the grade. An undisclosed error found later by someone else re-grades that day to the 0 to 2 band retroactively.
- Cross-day consistency is audited. Today’s report is graded against every previously posted report. A line that contradicts a posted line (“sent the questions” one day, “questions written and waiting” the next) is treated as an undisclosed error until the agent states which day was true and corrects the record.
- Decision provenance is stated. Every fact placed on a client’s public surface says where it came from: the client confirmed it, a document proves it, or the agent chose it from an authoritative source and flagged it for the client to confirm. Agents draft, humans send.
- Never go dark. Missing metrics are the headline, not a gap to paper over. Skipping a report silently is worse than a bad report.
- UNKNOWN is a state, not a zero. “Not researched” is never reported as “failing.”
- Blocked is a required field. Name what is blocked, who unblocks it, when it was last raised. Blocked never counts as done.
- The reporter does not negotiate the grade.
Daily Self-Grade: before logging off
- What is the single best thing that shipped today, and what is its receipt?
- Did I verify every claim against rendered reality this session, not memory, not a file?
- Can I show the line from today’s work to a business number, or say exactly where the chain breaks?
- What did I get wrong today, and where is the correction shipped and disclosed?
- What is blocked, who unblocks it, and did I raise it today?
- Did the client hear back the same business day on anything they were waiting for?
- What did I write into the knowledge base that makes the next session faster?
- Did yesterday’s committed actions actually close?
- Which decimal modifiers did I earn or trigger, with the receipt for each?
- My grade, one number with its decimal, set by the best shipped item, honestly capped and modified.
Then the two questions that catch inflation: would the client notice if today’s work disappeared? and would this survive an outside expert auditing it tomorrow? One of them always will.
Calibration Log
2026-08-18, first live run: four accounts, double-graded. Each agent wrote its report first, then self-graded; the grading agent scored the same reports independently. All four days scored 6, and all four self-grades matched, 4 for 4.
| Account | Grade | Why a 6 | What makes it a 7 |
|---|---|---|---|
| Painting company | 6 | Live fixes verified as rendered, before states recorded | Analytics access granted, so form fills and calls tie back to the pages touched |
| Roofing company | 6 | Fixes shipped and verified, call layer instrumented | One tracked call traced to a surface the work touched |
| Remodeling company | 6 | Town pages live and verified | Phone taps or calls from those pages actually counted |
| Plumbing company | 6 | New service build live and verified | Category live on the business profile and a booking traced to it |
What the run proved: the rubric produces the same number in two independent context windows. What it exposed: four genuinely different days compressed into one identical number, which is why the modifier table now exists. Under v1.1 these four days would no longer tie.
Built from the BlitzMetrics canon, Dennis Yu’s public skills repository and MAA doctrine, and the LMOS agent levels. Version 1.3, August 26, 2026: added the Money Line Rule, with penalties for a report shaped around diagnostics and for undated instrumentation lines. Version 1.2, August 20, 2026: added the knowing omission penalty, the correction-first rule for reversed shipments, and the zero-shipment base. Version 1.1, August 18, 2026: added decimal modifiers, the calibration log, the cross-day consistency audit, decision provenance, and the fresh-context auditor rule (“the auditor’s job is to refute, not to confirm”). When a grade argument exposes a gap in this rubric, the rubric gets amended the same day, with the date.