Factor 1000: Why a Vision Model Misreads Receipts
A bakeoff across 16 real receipts exposed a systematic factor-1000 error. Why a deterministic guard sits above the model, not inside it.
9 posts in this category.
A bakeoff across 16 real receipts exposed a systematic factor-1000 error. Why a deterministic guard sits above the model, not inside it.
One /discovery run, several parallel agents, a live-browser test: how I had my own website audited, and why a human still makes every final call.
Sven is my personal AI raven: he folds mailboxes, calendar, CI and server status into one Morning Pulse in Discord. He proposes, I decide and publish.
Three agents run around the clock: Sven scouts AI ideas, BuilderBob runs agenticbuilders.at with 19 cron jobs; session-orchestrator drives a gated coding loop.
An open-source multi-agent tool matures in 9 days across 18 sessions: from the /plan to /evolve lifecycle across four runtimes. What daily self-use made of it.
CLOUD Act, GDPR, self-hosting: which AI stack fits your project? A practical framework with real costs and three reference scenarios.
After 3,000+ sessions with Claude Code: what prevents AI agents from drifting in long-running projects. CLAUDE.md, Rules, Session Handovers and Quality Gates.
Three practical conclusions from the 2026 Gartner and HBR work-trend reports, viewed through my own work with AI agents.
What sits between an AI prototype and daily use: data, failure handling, ownership, and deciding when an experiment has answered its question.