DTWdailytechwire
Tech Intelligence, Wired Daily
Products

OpenAI's Bet on Workplace Agents Faces the Discovery Problem

Head of product Thibault Sottiaux oversees the company's push into enterprise, where "product of discovery" meets token economics and the question of whether white-collar workers are ready for autonomous AI.

AS
Arjun S. Mehta
AI Correspondent · Bengaluru
Aug 26, 2026
7 min read
OpenAI's Bet on Workplace Agents Faces the Discovery Problem
OpenAI's Bet on Workplace Agents Faces the Discovery ProblemCredit: OpenAI

From Code to Cubicles

OpenAI has spent two years teaching software engineers to trust autonomous agents. Now it wants the rest of the office to do the same. Thibault Sottiaux, who leads the company's core product organization covering API, ChatGPT, and the developer-focused Codex tool, is steering that expansion. The vehicle is ChatGPT Work, a platform designed to bring agentic capabilities to non-technical professionals. The company recently disclosed that the service crossed 20 million users, a milestone OpenAI frames as validation that "the world seems to be ready" for AI that can autonomously handle multi-step tasks like document synthesis, slide generation, and research.

At DailyTechWire, we've tracked enterprise AI rollouts across Asia and the U.S. for the past eighteen months. The pattern is consistent: adoption spikes when tools stay close to familiar workflows, then stalls when autonomy increases and transparency decreases. OpenAI's approach inverts that logic. Sottiaux describes the design philosophy as "getting out of the way of the model," minimizing interface chrome and letting the system decide how to execute requests. That stands in contrast to competitors like Anthropic, whose Claude for Work presents users with explicit choices and A/B paths during task execution.

The Diffusion Playbook

Sottiaux frames the transition from Codex to Work as a question of maturity and audience forgiveness. Developers tolerate rough edges, incomplete outputs, and the occasional hallucination because they can audit and fix code. Knowledge workers operating in natural language have less room for error and fewer tools to verify output. OpenAI's answer is to delay broader release until models reach what the company considers a "level of maturity" sufficient for safety and delight.

That maturity threshold is subjective. GPT-5.6, the model powering current Work deployments, shows improved performance on document processing and report generation compared to earlier versions. But the platform still lacks the structured decision trees that enterprise software buyers expect. Sottiaux acknowledges this tension, describing the product strategy as "discovery" rather than roadmap execution. The team builds capabilities, observes how users stretch them, then doubles down on the most successful patterns.

This iterative, deployment-first model borrows from consumer social products rather than enterprise SaaS. It assumes that real-world usage will surface both capabilities and failure modes faster than internal testing. For OpenAI, which has staked its reputation on responsible AI development, that trade-off carries reputational risk. The company's safety benchmarks show strong alignment scores, but alignment in controlled settings does not always predict behavior in messy, multi-stakeholder workflows.

The Token Economics Nobody Wants to Discuss

One unresolved tension in the Work model is cost predictability. OpenAI charges a flat $20 per month for Plus subscribers who access Work features, but token consumption varies wildly depending on task complexity. Users who lean on the platform for heavy document analysis or iterative research can burn through token budgets that far exceed the subscription price. Sottiaux sidesteps the question, pointing instead to the 80 percent price reduction OpenAI announced with its Luna infrastructure update. He frames efficiency gains as a permanent deflationary trend, suggesting that today's frontier capabilities will become cheaper over time.

That logic holds for training and inference hardware, where Moore's Law analogues still apply. It breaks down when models grow more capable and users demand more from them. A professional who discovers that ChatGPT Work can draft a board presentation will ask it to draft ten. Usage expands to fill available capability, a dynamic familiar to anyone who has watched cloud bills scale with adoption. CFOs evaluating enterprise AI deals are starting to ask hard questions about per-seat costs versus per-token exposure, especially in markets like Singapore and Seoul where finance teams have deep SaaS procurement experience.

OpenAI's answer, for now, is to absorb the gap and bet that volume and efficiency improvements will close it before margin pressure becomes unsustainable. That works in a growth-at-any-cost environment. It becomes harder to defend if revenue per user stagnates while compute costs remain high.

The Permission Problem

Another friction point is data access. Work's value proposition depends on integrating with email, messaging platforms, internal documents, and collaboration tools. OpenAI rolled out iMessage integration in August 2026, adding to existing Gmail and Slack connectors. Each new integration expands the model's context window and, in theory, its usefulness. It also expands the attack surface and the trust burden.

Sottiaux's response to privacy concerns is to emphasize model alignment and safety benchmarks. OpenAI publishes evaluation scores for adversarial prompt resistance, data leakage prevention, and output filtering. Those scores matter, but they address technical risk, not organizational risk. A model that never leaks data can still summarize a sensitive email thread in a way that violates internal policy or creates compliance exposure. Enterprise buyers in regulated industries want contractual guarantees, audit trails, and the ability to sandbox deployments by department. OpenAI has built some of that infrastructure for API customers, but the consumer-grade Work experience offers less granular control.

This gap matters more in Asia than in the U.S. Companies in Japan and South Korea have stricter data residency requirements and more conservative approaches to third-party cloud services. OpenAI has not disclosed where Work inference runs or whether it offers regional data isolation. Those details will determine whether the product can move beyond early adopters in those markets.

Discovery as Strategy, or as Excuse

Sottiaux describes product development as "discovery," a term that captures both the excitement and the risk of building on top of rapidly improving models. When capabilities jump between versions, product teams must decide whether to expose new features immediately or hold them back until the user experience catches up. OpenAI has consistently chosen speed, shipping voice interfaces, multimodal input, and agentic workflows as soon as the underlying models support them.

That approach works when the user base is technical, curious, and tolerant of failure. It becomes more complicated when the product moves into regulated workplaces where mistakes have consequences. A coding agent that writes buggy Python is annoying. A work agent that misinterprets a contract clause or generates a factually incorrect client memo is a liability.

OpenAI's iterative deployment philosophy assumes that broad usage will surface problems faster than internal QA, and that the company can patch issues quickly enough to prevent lasting harm. That assumption has held so far, but it scales poorly. As Work moves deeper into enterprise accounts, the cost of a high-profile failure increases. A single incident involving data leakage, a discriminatory output, or a costly error could stall adoption across entire sectors.

The Simplicity Trap

Sottiaux repeatedly emphasizes simplicity: minimal UI, natural interaction, no learning curve. The goal is to make AI feel like a conversation rather than a tool. Voice input, which OpenAI introduced earlier this year and which has seen strong uptake, pushes that vision further. The less users have to think about prompts, parameters, or workflows, the theory goes, the more they will use the product.

But simplicity in interface design can obscure complexity in execution. When a user asks ChatGPT Work to "prepare a summary of this quarter's performance," the system makes dozens of decisions: which documents to prioritize, how to weight qualitative versus quantitative data, what tone and length to target, whether to flag anomalies or smooth them into a narrative. All of those choices are invisible to the user, who sees only the final output. If the output is wrong, or incomplete, or subtly misleading, the user has no way to trace the reasoning.

Anthropic's approach with Claude for Work, which surfaces decision points and asks users to choose between paths, sacrifices speed for transparency. OpenAI's approach bets that most users prefer speed and will tolerate occasional errors in exchange for convenience. That bet may prove correct for individual contributors working on low-stakes tasks. It is less clear whether it holds for managers, executives, or teams working under regulatory scrutiny.

What Comes After the Dopamine Hit

Sottiaux mentions that Codex users know him as the person who raises token limits when the product hits growth milestones. It is a small detail, but a telling one. OpenAI has built a user base conditioned to expect expanding capability and falling costs, a dynamic that mirrors the trajectory of consumer internet services in the 2010s. The question is whether that dynamic can sustain a business model built on compute-intensive inference rather than ad-supported content.

The company's Luna pricing announcement, which cut API costs by 80 percent, signals that OpenAI believes efficiency gains will outpace capability growth. If that holds, Work can remain a flat-rate subscription product even as users demand more from it. If it does not, OpenAI will face a choice: raise prices, impose usage caps, or continue subsidizing heavy users at a loss.

For now, the company is betting on volume. Twenty million Work users, even at $20 per month, represents $400 million in annualized revenue, a meaningful line item but not yet a business that justifies OpenAI's valuation. The real prize is converting those users into enterprise accounts, where per-seat pricing can reach hundreds of dollars per month and where organizations commit to multi-year contracts. That transition depends on solving the trust, transparency, and cost predictability problems that Sottiaux's "discovery" model has so far deferred.

OpenAI has proven that it can build models that impress early adopters and generate viral growth. The next test is whether it can build products that CFOs, compliance officers, and risk managers are willing to bet their operations on. Discovery works well in the lab. In the boardroom, it looks like a lack of planning.

Read next
Products

Amazon Pushes Device Costs Up by 60 Percent Amid Memory Crunch

Arjun S. Mehta · 4 min
Products

When Paying for AI Stops Making Sense

Priya Nair · 6 min
Products

Lenovo Responds to Legion Go Failures as Repair Costs Mount

Marcus Halloran · 5 min
Spot something wrong? Email corrections@dailytechwire.com. We log every correction publicly.