I open-source the tooling I wanted while building that:
tokenprof: profiles what is actually consuming your context window, per turn, per tool. https://github.com/muhammadwaqar12/tokenprof
awesome-agent-failures: production failure modes for LLM agents: symptoms, reproductions, mitigations. https://github.com/muhammadwaqar12/awesome-agent-failures
Mostly here for context engineering, agent evaluation and trajectory replay, and where LLMs quietly break on procedural knowledge.
Open to AI engineering and forward-deployed roles. m_waqar@live.com - github.com/muhammadwaqar12