Mostly about building products, healthcare tech, and lessons learned along the way.
1 article on llm evals.
Most teams skip real evals and wonder why their AI products degrade in production. The framework that holds up: from 30-minute manual reviews to binary scoring to knowing when your eval suite is finally doing its job.