On 8 September 2026, OpenAI published a striking customer result: 1Password estimated a 553% return on investment, or ROI, from Codex. Crucially, that figure values additional engineering capacity; it does not establish cash savings. OpenAI’s 1Password case study

For a German SME deciding whether to expand an AI pilot, that distinction changes the budget conversation. Faster replies or reports do not automatically reduce payroll. Their value depends on whether employees can clear a backlog, improve service or take on useful work. Before buying more licences, managers need to connect the time saved to a business outcome and compare that benefit with the complete cost.

Read the assumptions behind the return

1Password’s model values annual capacity at $783,750 for 50 active developers. Its inputs include a reported 20.9% productivity improvement, $250,000 annual cost per developer, 40% attribution to Codex and 75% capacity realisation. The measurement window and productivity definition are unspecified. Case study methodology

In the same publication, Nancy Wang, 1Password’s chief technology officer, describes faster progress from planning to production. This is a customer’s perspective in its supplier’s marketing, rather than independent validation. Wang’s published comments

The useful lesson is to expose the assumptions. How much improvement came from AI? How much usable time remained after checking its work? What happened to that time?

Research supports testing at the level of the task

Independent research gives reasons for both interest and caution. In Generative AI at Work, Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,172 support agents associated with one software company, with rollout concentrated in autumn 2020 and winter 2021. AI assistance increased issues resolved per hour by about 15% on average. Less experienced workers benefited more; the most skilled saw small declines in quality. This was a particular support workflow, not a forecast for today’s tools or European SMEs. Original research

METR’s experiment using early-2025 tools reached a different result: 16 experienced open-source developers took 19% longer across 246 tasks when AI was allowed. That narrow setting does not establish that AI generally slows work. In February 2026, METR reported that follow-up research suggested improvement, but participant selection and time-measurement problems prevented a reliable estimate of the current effect. Original experimentfollow-up assessment

Joel Becker, an evaluation researcher and author of METR’s 11 May 2026 analysis, argues that faster work and more valuable work need separate measurement. He also cautions that reported productivity gains can overstate reality. This is a published methodological assessment from a research nonprofit, not a supplier’s customer testimonial. Becker’s analysis

Taken together, these findings support testing the intended workflow, including its users and quality requirements, before borrowing anyone else’s productivity percentage.

Decide what the saved time is for

Before buying more licences, identify the intended benefit. Reduced overtime or external spending can produce cash savings. Faster handling of an existing backlog can create operational value. Additional profitable orders can contribute to earnings, but only if demand exists and the rest of the business can fulfil them.

Keep those benefits separate. Do not value the same hour once as labour savings and again as additional output. Where employees regain breathing room, record that benefit honestly and decide whether it justifies the expense. More manageable workloads can be a valid objective without being described as a reduction in payroll.

For an ROI calculation, subtract the total project cost from the benefit value, divide by that cost and multiply by 100 for a percentage. Use the same period for both. Label the result according to what the benefits contain: cash savings, estimated capacity value or a mixture. A precise percentage cannot repair an uncertain valuation.

Run a pilot over 90 days

For a frequent, measurable workflow, such as preparing standard customer responses or checking supplier documents, a 90-day pilot provides a practical structure. Treat the timetable as a starting point; the volume and variety of work determine whether the evidence is sufficient.

Days 1–15: establish the baseline. Select one process and name its owner. Record completed work, total handling time, error rates and rework. Define an acceptable result before introducing AI. For a German customer-service team, include the accuracy of German-language replies and exceptions requiring specialist judgement.

Days 16–60: compare like with like. Where practical, randomly assign comparable cases to AI-assisted and existing workflows. Balance case difficulty and staff experience. Include training time, prompting, review, corrections and work transferred to colleagues. Record the model and configuration so that a product change does not silently alter the comparison.

Ask users what helps, but pair their answers with process records. Track whether faster drafting also means faster resolution, and whether customers return with the same unresolved problem.

Days 61–90: test the business case. Count subscriptions, usage charges, integration, supervision and maintenance. Separate setup spending from recurring costs. Agree how to value usable capacity and test a conservative scenario with lower adoption, more review or fewer eligible cases.

Set the decision rule before seeing the result. Expand only if the agreed quality threshold holds and the benefit clears the business’s cost hurdle. Extend the pilot when case volumes are too low to support a decision. Stop or redesign it when review effort absorbs the saving.

Put the numbers through a reality check

Consider a hypothetical ten-person operations team. Each employee saves 20 minutes per working day after review, across 20 days a month. That releases roughly 67 hours. At an assumed employment cost of €40 an hour, the initial capacity valuation is about €2,667 monthly.

Suppose only half those hours can be redirected to useful work. The modelled value falls to approximately €1,333. With €600 in recurring monthly costs, the remaining balance is about €733 before setup costs. These are illustrative assumptions, not a forecast or measured customer result.

That €733 is not automatically profit. The manager must show what the redeployed hours accomplish. If the value depends on extra sales, use the contribution after the associated costs, rather than sales revenue alone.

This approach suits work with repeatable cases and observable outcomes. It is less conclusive for rare strategic decisions or long research projects, where quality and downstream value take longer to assess.

The next step is small: choose one workflow, write down the baseline and ask its owner what will happen to any hours released. An AI investment becomes easier to defend when that answer survives comparison with the costs and the finished work.



Leave a Reply

Your email address will not be published. Required fields are marked *