Rubrics for evaluating AI agents
A repeatable, rubric-based way to evaluate how AI agents behave when they act on real systems through MCP tools β consistent scores instead of one-off manual checks.
Hi, I'm Nico π
Over a decade in the IT industry, currently focused on how AI agents behave when they operate real systems. I write about technology and life in Japan, and I help engineers and teams through mentoring and consulting.
I have recently revisited a topic that fascinated me several years ago: binaural waves (also known as brain waves or binaural beats).
Welcome back! Today I want to share a story about how I nearly fell for a piece of software. Well, sort of.
Hello there! π I'm excited to share some of the interesting things Iβve been cooking up recently. As part of my job, I focus on creating innovative solutions and fillingβ¦
A repeatable, rubric-based way to evaluate how AI agents behave when they act on real systems through MCP tools β consistent scores instead of one-off manual checks.
Using Distrobox containers on an immutable Linux desktop to keep every toolchain isolated, disposable and reproducible.
This site: a Jekyll landing page plus a Chirpy blog, built in one pipeline and served from Cloudflare's edge, mirrored on GitHub Pages.
One-on-one guidance for engineers growing their skills and career.
Design, test and evaluate AI agents that act on real systems.
Faster, reproducible development environments for your team.