ER10 · July 28, 2026 · 1Cifer
Frontier AI models still struggle with real work tasks — worth knowing before you deploy
Researchers tested frontier AI models on tasks close to real office work — and the results are sobering: where company context, unstated conditions and internal rules matter, models err noticeably more often than on academic tests. Pretty benchmark scores still convert poorly into reliable "works like an employee" performance.
The gap is explainable: work tasks rarely arrive as a clean problem statement. People lean on what isn't in the request — the client's history, tacit agreements, knowing "how we do things here". Without that context a model produces a formally correct but useless answer.
For business this is not a reason to postpone AI but a usage manual. Two rules work: narrow the task and supply context. That is exactly why an agent living inside company data beats a universal chat: in 1Cifer every company, department and project has its own workspace, and the agent answers not "on average across the internet" but within a specific circuit with its numbers and rules. And decisions with a high cost of error should stay with a human for final review in any case.


