Not every agent step deserves the largest model
What changed after treating agent traces as routing signals instead of passive logs.
I learn by understanding systems from first principles, building things to test those ideas, and seeing what happens when they meet reality.
I'm trying to understand how LLM inference actually works.
What changed after treating agent traces as routing signals instead of passive logs.
Why failure taxonomy became the foundation of useful evaluation and root-cause analysis.
What a cross-component test suite revealed that isolated unit tests could not.