OpenAI is making a simple business argument: stop judging AI by how many people bought access or how cheap the tokens are, and start judging it by work accomplished. The company’s new framing — “useful intelligence per dollar” — says the real test is whether AI completes valuable tasks, at acceptable quality, at a cost that makes sense once you include retries, human review, and employee time. That matters for Canadian businesses right now because CFOs, operations leaders, and IT teams are already trying to decide which AI tools deserve a budget line and which ones are just expensive demos.
The pitch is not just about model price. It is about outcome economics. A model that is cheaper per token can still be costly if staff have to babysit it. A pricier model can be the better buy if it gets the job done in one pass. That is a useful idea for Canadian firms in finance, logistics, construction, healthcare, and professional services, where the cost of a bad answer is often bigger than the cost of the software itself. But the same framing also risks giving vendors a new way to sell ambition before proof. That is where the debate starts.
Alex: the case for optimism
Kevin is right to be suspicious of vendor language, but I think he is missing the bigger shift here. OpenAI is at least asking the right question: what work actually got done? That is a much healthier standard than counting logins or seats. Canadian businesses have been burned before by software that looked busy but did not change the day-to-day grind. If AI is going to earn its keep, it should save time on real work — the forecast prep, the customer follow-up, the document review, the internal reporting that eats up entire afternoons.
That is why this “useful intelligence per dollar” idea matters. It pushes leaders to measure outcomes instead of vibes. In practice, that could help a Calgary accounting firm compare tools based on how many files get reconciled correctly, or help a Prairie manufacturer see whether AI is actually reducing scheduling friction. It also gives smaller Canadian firms a way to compete with larger ones. If a five-person team can use AI to draft, sort, summarize, and triage like a much bigger shop, that is not hype — that is capacity.
And yes, I know the fear: “If AI does more work, what happens to the workers?” Kevin keeps raising that, and it is a fair concern. But this scorecard does not have to mean fewer people. It can mean better use of people. A finance team that spends less time moving numbers between tabs can spend more time explaining what the numbers mean. A legal team that gets a first-pass contract review from AI can focus on judgment, negotiation, and risk. That is not a fantasy. It is exactly how automation has usually improved work: by removing the dullest parts so humans can do the parts that matter.
OpenAI also makes an important point about dependability. If AI is going to move from drafting to acting, it has to be accurate, consistent, and escalated properly. That is not a reason to avoid it; it is a reason to use it carefully. Canadian companies already manage risk in payroll, invoicing, and compliance systems. They can do the same with AI. Start with one workflow. Define “done.” Measure success. That is practical, not reckless.
Kevin’s caution is useful, but his framing can freeze people in place. If every AI tool is treated as a potential trap, businesses will keep paying humans to do work software could already handle well enough. That is a hidden cost too. The better path is to test, measure, and improve. This is exactly the kind of thinking we cover in more of our coverage — not because AI is magic, but because Canadian firms need a way to separate useful tools from shiny distractions.
OpenAI’s scorecard is not perfect. No vendor scorecard is. But it nudges the market toward something healthier than token-count theatre. For Canadian business owners, that is a win: fewer vanity metrics, more actual work completed, and a clearer case for adopting AI now instead of waiting for the perfect moment that never comes.
Kevin: the case for caution
Alex is too quick to treat a new metric as if it solves the underlying problem. It does not. “Useful intelligence per dollar” sounds disciplined, but it is still a vendor-defined scorecard from a company selling the thing being measured. That should make every Canadian buyer nervous. If the seller gets to define usefulness, cost, and success, the numbers can look better than the reality on the ground.
The first issue is that OpenAI’s framing quietly shifts attention away from the hard part: trust. A task can be “completed” and still be wrong, incomplete, biased, or unsafe. A contract review that misses one bad clause is not useful. A support reply that sounds polished but gives the wrong policy answer is not useful. A finance workflow that saves time but introduces a hidden error is not useful. The company says dependability matters, which is true, but real dependability is not a slogan. It is testing, monitoring, escalation rules, audit trails, and people who know how to catch failures before they become expensive.
That is where Canadian businesses can get hurt. A small firm in Calgary may not have a dedicated AI governance team. A mid-sized manufacturer may not have the staff to review every edge case. A healthcare or finance organization cannot afford “mostly right.” Yet the pressure to cut costs will push leaders to automate faster than they can supervise. Once managers start chasing “useful intelligence per dollar,” the temptation is to define useful too loosely and let the system run ahead of the controls.
And let’s talk about jobs, since Alex keeps trying to soften that concern. Yes, automation can remove drudgery. It can also remove entry-level work, support roles, and the repetitive tasks that juniors use to learn the business. If AI takes over the first pass on research, reporting, document review, and customer triage, what exactly is left for the next generation to practice on? Canadian workers are not wrong to worry that “amplifying people” is often how companies describe headcount reduction before they say the quiet part out loud.
There is also the vendor lock-in problem. If a company rebuilds its workflows around one AI provider’s models, tools, and measurement language, switching later becomes expensive. That matters in Canada, where many businesses are already cautious about cloud concentration, data handling, and cross-border dependency. A scorecard that looks clean on a slide can hide long-term operational risk.
OpenAI’s own examples make my point. The company says a more capable model may cost more per token but less per successful task. Fine. But who decides what counts as a successful task? Who checks the failures that never show up in the dashboard? Who carries the liability when the AI gets it wrong? Those questions matter more than the new label. Alex wants firms to “test and improve.” I agree. But most firms do not have the discipline to do that well, and the ones that do will still need to assume the model is fallible by default.
So no, this is not a breakthrough in AI economics. It is a reminder that the sales pitch has evolved. The risk for Canadian businesses is that they will mistake a better metric for a safer product. It is not safer. It is just easier to justify.
Donald: the balanced read
Both of my colleagues are reacting to the same thing, but they are emphasizing different parts of it. OpenAI is trying to move the market from input metrics to outcome metrics. That is a real shift. For years, software buyers have been asked to judge systems by adoption, usage, or raw cost. AI complicates that because the cheapest output is not always the cheapest successful outcome. Once you include retries, human review, and workflow delays, a model with a higher per-token price can sometimes be the better business choice.
That is the strongest part of the argument. It is also the least controversial. Canadian businesses already understand the difference between a low sticker price and a low total cost of ownership. The same logic applies here. If AI helps a team complete more work with fewer handoffs, that has value. If it only creates more drafts for humans to clean up, the savings disappear.
Kevin is also right that the new scorecard does not remove the need for governance. In fact, it may increase it. The more AI moves from drafting to acting, the more important it becomes to define acceptable error rates, review thresholds, escalation paths, and accountability. That is especially true in Canadian sectors where mistakes are costly or regulated: finance, healthcare, legal services, insurance, and parts of the public sector. A metric like useful intelligence per dollar is only meaningful if the “useful” part is measured honestly.
Where I part ways with Kevin is on the idea that this is mostly a repackaged sales pitch. It is a sales pitch, yes. But it is also a useful prompt for buyers. Many Canadian firms still ask the wrong question: “How many users will we buy?” or “How cheap is the model?” OpenAI is pushing them to ask: “What work will this actually complete, and what will that save or unlock?” That is a better conversation for procurement, finance, and operations teams.
Where I part ways with Alex is on speed. He is right that businesses should test and learn, but he can sound too relaxed about the operational burden. Most firms will not get this right on the first try. They will need narrow pilots, clear metrics, and a willingness to stop projects that do not prove themselves. In other words: optimism, but with paperwork.
The Canadian angle is straightforward. Labour is tight in many sectors. Productivity growth remains a real concern. AI tools that genuinely reduce admin work, speed up analysis, and improve turnaround times will be attractive. But any business that adopts them without controls risks trading labour cost for error cost. The scorecard is useful only if it is applied to real workflows, with real consequences, in real Canadian conditions.
What this means for Canadian businesses
The practical takeaway is not “buy AI” or “avoid AI.” It is to stop evaluating AI like ordinary software and start evaluating it like a work system. Canadian businesses should pick one workflow, define success in plain language, track the full cost of getting it right, and measure how often the system needs human correction. If the tool saves time but creates risk, the savings may be fake. If it improves throughput without lowering quality, it may be worth expanding.
For smaller firms, that means resisting the urge to automate everything at once. For larger organizations, it means building governance into procurement, not bolting it on later. And for workers, it means being honest that some tasks will shrink or disappear while others become more valuable. That is the real tension in the AI age: not whether AI can do work, but which work should be delegated, which should be supervised, and which should stay human.
OpenAI’s scorecard is a useful provocation because it forces that conversation. It does not settle the debate. It sharpens it. For Canadian businesses, that may be the most valuable thing about it.
If you want more plain-English analysis of how AI is changing Canadian business, keep following davision.ca.
