Blog
What companies look like when agents do the execution: the evidence, the counter-evidence, and where we think it lands. Every figure traced to a source we actually read. Start with why an LLM judge needs its own controls, the method behind how this series treats model-scored evidence.
LLM judges can approximate human preferences, but Nazr saw a score move 0.33 when the generator's model family joined the panel. Here are the controls.
Nazr's benchmarks found that model size mattered only in combination with better input and whole-plan context. Here is what each test established in practice.
Measure an AI plan as a whole tree, not polished branches. Nazr's six-part rubric exposed duplication, conflict, uneven scope, and weak sequencing in practice.
Nazr ran a nine-tree paired test of its plan critic. The composite stayed 3.31 to 3.31, revealing what the critic changed and what judges could actually see.
AI benchmarks fail when the metric, execution path, judge, or sample differs from the production decision. Nazr found each failure directly in its own evals.
Evaluate an AI planning system across capture, model, architecture, judges, whole-plan coherence, decision policy, and the boundary for safe human review.
Lovable reported $100M ARR with 45 employees, Gamma with about 50, Cursor with roughly 60. The verified figures, the pre-AI baseline, and what happens next.
Meta, Amazon, Google, and UPS removed management layers on the record, yet mid-size companies hold a steady share of US employment. The verified record.
No measurable mass displacement yet, and Klarna and Duolingo walked back their AI replacement claims. The verified signal is a 19% entry-level hiring gap.