# IBM Is Certifying Thousands of Consultants on OpenAI. The Models Cost $20 a Seat.

IBM just embedded GPT-5.6 into a consulting arm that earns 11.7 cents of segment profit per revenue dollar. The same models rent for $20 a seat. Here is what the gap buys.

Author: J.A. Watte
Published: August 21, 2026
Source: https://jwatte.com/blog/ibm-openai-what-large-firms-are-buying/

---

On 13 August 2026, IBM announced a partnership with OpenAI. The release says IBM will build "industry-specific solutions for financial services, government, telecommunications, and retail, as well as key enterprise domains such as finance, procurement, customer operations, and HR." It embeds OpenAI's GPT-5.6 models, Codex and ChatGPT Work into IBM Consulting Advantage. And it stands up a new unit: "IBM will also launch a dedicated OpenAI Practice, with thousands of consultants and engineers obtaining expert-level certifications under the OpenAI Partner Network."

I read that, then I opened OpenAI's own pricing pages. GPT-5.6 Sol, the flagship model in the deal, costs $5.00 per million input tokens and $30.00 per million output tokens on the standard API tier for short-context requests as of 21 August 2026. Long-context requests are $10.00 and $45.00. A ChatGPT Business seat, which now includes Codex, is $25 per user per month billed monthly or $20 billed annually, with a two-seat minimum.

Same models. So the question worth asking is not "what does IBM know that I don't." It is: what is the gap between $20 a seat and an IBM engagement actually buying?

The answer is not intelligence. It is integration into decades of fragmented systems, plus certified bodies to do that integration, plus a counterparty large enough to sue. A two-person company has almost none of that fragmentation. That absence is the single biggest structural advantage a small business has right now, and nobody is going to tell you about it because there is no fee attached to it. It is also mostly untaken. Census survey data from December 2025 to May 2026 puts overall business AI use between 17 and 20 percent, with under 20 percent of firms with four or fewer employees using AI at all, and I could not find a single named business under 100 employees with a measured, primary-sourced result. Every case study below is a large company, because those are the only ones that publish.


## IBM Consulting is the smaller, flatter, thinner half of the company

The OpenAI deal runs through IBM Consulting. Worth knowing what that is.

In the quarter ended June 2026, reported 22 July 2026, IBM Consulting brought in $5.3 billion, flat year over year and up 1 percent at constant currency. IBM Software brought in $7.8 billion, up 5 percent. IBM Infrastructure brought in $3.8 billion, down 7 percent. Total company revenue was $17.2 billion, up 1 percent.

For the full year 2025, Consulting revenue was $21,055 million, up 1.8 percent as reported and 0.4 percent adjusted for currency. Segment profit was $2,464 million on an 11.7 percent segment profit margin, which IBM notes was up 1.8 points on the prior year. Read that margin again. The entire consulting arm converts a bit under twelve cents of every revenue dollar into segment profit. That is a business built on utilisation of people, and it prices accordingly.

For scale, IBM's total 2025 revenue was $67.5 billion with $14.7 billion of free cash flow, and Software was roughly 45 percent of revenue. So the OpenAI partnership is being run through the smaller, lower-margin, flat-growth half of IBM, not the half that is actually growing.

One thing I could not establish, and I looked properly: IBM does not publish a headcount. Not for the company and not for Consulting. The FY2025 10-K defers human capital disclosure to the annual report, and the annual report's only staffing number is that over 200,000 employees participated in an engagement survey in 2025, which is a participation count and a floor, not a total. Plenty of directory sites repeat a figure for "IBM consultants." None of them sources it to IBM. So when the press release says "thousands of consultants and engineers" will be certified, there is no published denominator to compare it against.

Accenture, by contrast, states its headcount plainly: approximately 799,000 people. Accenture's Consulting revenue alone in its quarter ended 31 May 2026 was $9.3 billion, which is 75 percent larger than IBM's entire Consulting segment in the calendar quarter ending a month later. Accenture's total revenue for the quarter was $18.72 billion, up 6 percent in US dollars, on new bookings of $19.32 billion, which were down 2 percent.

## The AI metric everyone quotes counts orders, not outcomes

You will see IBM's "generative AI book of business" cited constantly. It is worth knowing exactly what it counts, because IBM tells you and almost nobody repeats it.

From IBM's own 8-K exhibit filed 22 July 2026: "It is calculated as inception to date Software transactional revenue, plus new SaaS Annual Contract Value and Consulting signings related to specific offerings. Since second-quarter 2023, approximately one-fifth of this book of business comes from Software, and the remaining four-fifths is Consulting."

Four-fifths of it is Consulting signings. A signing is an order. It is not delivered work, not recognised revenue, and not a measured client outcome. It is a contract someone signed. The last dollar figure IBM published for it is "more than $12.5 billion since inception to date," in the chairman's letter of the FY2025 annual report published 24 February 2026.

You may also see $9.5 billion quoted for the same metric. Both numbers are real IBM figures from different quarters. The $12.5 billion is the year-end 2025 figure and is the current one. If you see the smaller number without a date attached, it is stale.

Now the part that made me sit up. In the second quarter of 2026, IBM stopped attaching a dollar figure to that metric at all. The definition is still in the 8-K boilerplate. The number is gone. What the CFO said instead was: "Generative AI represented about 50% of our signings in the quarter and now makes up over 30% of our backlog."

Percentages of signings, not cumulative dollars. And percentages need a denominator. Here is IBM's: Consulting signings fell 13.3 percent in 2025 to $21,757 million, with a trailing-twelve-month book-to-bill of 1.03 and year-end backlog of $31.9 billion. A book-to-bill of 1.03 means the order book is barely more than replacing the revenue being burned off it.

So "generative AI is half of Consulting signings" and "Consulting signings fell 13 percent last year" are both true at once. Signings did return to growth in 2026, up 6 percent in the June quarter, which IBM called its second consecutive quarter of growth. That growth is real, and it is off a base that dropped by an eighth. Meanwhile Consulting revenue grew 1 percent.

Accenture did the same thing in the same year. Its last published generative AI figures are for fiscal 2025: $5.9 billion of generative AI new bookings for the year against $80.6 billion of total new bookings, and $2.7 billion of revenue from generative and increasingly agentic AI against $69.7 billion of total revenue. That is roughly 3.9 percent of revenue. Accenture also renamed the category from "Gen AI" to "advanced AI" in September 2025. Its June 2026 quarterly release and earnings call carry no advanced-AI bookings or revenue figure at all, only qualitative statements.

Two of the largest AI services sellers on earth both quietly stopped publishing their headline AI dollar metric in 2026, at exactly the moment the market got loudest about AI services. Neither announced that they had stopped. I would not read that as scandal. I would read it as a strong hint that the metric had stopped flattering them, and I would stop treating either number as a measure of value delivered.

To be fair to Accenture, it publishes staffing detail IBM does not: 40,000 AI and data professionals in FY23 rising to 77,000, over 550,000 staff trained in generative AI fundamentals, and more than 6,000 advanced AI projects delivered in FY2025.

## What an IBM consultant costs, from a price list anyone can read

The most useful pricing document in this whole story is not a press release. It is IBM's federal schedule.

IBM holds GSA Multiple Award Schedule contract GS-35F-110DA, with a current option period ending 20 December 2030. Its published Appendix C rate card for IT services sets a 2026 ceiling price of $401.05 per hour for a Consultant V, defined as a bachelor's degree and 12 years of minimum experience. An entry-level Consultant I on the same card, bachelor's and one year, is $248.34 per hour. Consultant III is $309.62. Architect V, 12 years, is $386.75.

These are ceiling rates including the Industrial Funding Fee, so a real order is at or below them. They are still the clearest public signal of what this labour is priced at.

Do the multiplication yourself. Five people at roughly Consultant III money, $309.62 an hour, 2,000 hours each in a year, comes to about $3.1 million at list. That is one modest team for one year. It is why "an enterprise AI consulting engagement" and "eight figures" are compatible statements once you add duration and headcount, without anyone being greedy about it.

There is one oddity in that document worth a paragraph, because it tells you something about how services are priced generally. IBM maintains a second GSA rate card for management consulting, and on that card a "Consultant I" with a bachelor's and one year of experience is $109.46 per hour in 2026. Same nominal job title, 2.3 times the price on the IT-services card. The IT-services card also has no education differentiation whatsoever; every category is listed as "Bachelors" and seniority is priced purely on minimum years. What you are buying is a contract vehicle and a labour category label, not a measured competence.

Set the two numbers next to each other. Two ChatGPT Business seats billed annually cost $480 for a year. One hour of a Consultant V at ceiling is $401.05. Your entire annual model bill for a two-person shop is roughly seventy-two minutes of one senior consultant at list price.

That comparison is unfair, deliberately, and the rest of this piece is about why.

## What OpenAI is building on its side

OpenAI launched its Partner Network on 14 June 2026. The structure, in OpenAI's words: "Partners can progress through three tiers: Select, Advanced, and Elite, each with a high bar for sales performance, technical capability, co-sell engagement, and deployment experience." Beyond tiers, partners can earn specializations "in high-impact areas such as Codex, cybersecurity, and agents."

The investment figure is $150 million, and the stated goal is to "train and enable 300,000 certified consultants by the end of 2026."

Three hundred thousand certified consultants. Against Accenture's 799,000 employees and IBM's undisclosed count, that is a serious channel build, and it is the actual product here. OpenAI is not selling IBM intelligence. It is renting IBM's client relationships and buying trained hands.

There is also a forward-deployed layer, which is the part most people miss. OpenAI is "piloting a Forward Deployed Experts program with a set of founding partners" designed to align partner practitioners with OpenAI's own Forward Deployed Engineering teams. Separately, on 11 May 2026 OpenAI launched the OpenAI Deployment Company to embed Forward Deployed Engineers inside customer organisations, with more than $4 billion of initial investment, 19 partner investment and consulting firms led by TPG, and an agreement to acquire Tomoro that brings roughly 150 experienced FDEs and deployment specialists from day one. OpenAI describes the model as starting from one customer's specific problem: "FDE teams work directly with customers to solve a specific problem, validate impact, and then identify patterns that can scale."

That is not a new idea. Palantir invented the role and named it in public in 2020: a Forward Deployed Software Engineer, internally a "Delta," is "a software engineer who embeds directly with our customers to configure Palantir's existing software platforms." Palantir's framing of the split is the sentence everyone paraphrases: a traditional engineer "focuses on creating a single capability that can be used for many customers," while forward-deployed engineers "focus on enabling many capabilities for a single customer."

Read that as a pricing statement and it explains the whole enterprise services market. Many capabilities for a single customer is expensive per customer by construction. It does not amortise. That is precisely the cost a company with one customer profile and one set of systems does not have to bear.

Two honest caveats about the OpenAI side.

**The tier claim is asymmetric.** IBM's release says IBM joins OpenAI's Elite tier. I could not find OpenAI publishing that, or any tier label for any partner. OpenAI names Select, Advanced and Elite in a single blog paragraph on 14 June 2026 and then never mentions tiers again. Its partner locator renders partner cards without tier badges, and IBM's own locator entry lists only "Countries served Global | Industry Cross-industry | Program partners AWS, Daybreak, and Oracle." No thresholds published, no partner tier assignments published. So "Elite tier" is currently a claim by the partner about itself, in a programme whose criteria are not public. That is not an accusation. It is a reason not to treat the label as third-party verification.

**"ChatGPT Work" is not the enterprise plan.** IBM's release names it as one of the OpenAI products being embedded, and it is easy to read that as the name of the business subscription. It is not. ChatGPT Work is an agent inside ChatGPT, launched 9 July 2026, that "can gather information across your apps and workflows to create finished materials like sheets, slides, docs, and web apps." It rolled out to Pro, Enterprise and Edu users first, then Plus and Business. The commercial plans are still ChatGPT Business and ChatGPT Enterprise. OpenAI separately uses "ChatGPT for Work" as an umbrella seat-count metric, which is a genuinely confusing collision of names.

## The four domains, checked against the evidence

IBM named four: finance, procurement, customer operations, HR. I went looking for defensible published outcome numbers in each. The results are uneven in a way that is itself informative.

**Finance.** BlackLine announced general availability of a multi-agent reconciliation system on 27 July 2026, reporting that "early adopters have achieved up to a 92% reduction in manual reconciliation preparation time." Print the qualifiers with the number: "up to," and "early adopters," meaning a self-selected group of the customers most motivated to make it work. It is a vendor claim about its own product. It is still the most concrete finance number I found.

**Procurement.** This is the thin one, and I am going to say so plainly, because it is one of the four IBM chose to name. The only usable primary source I found is Coupa's quarterly release of 15 December 2025, which says customers "used Coupa to manage over $425B of business spend, and realized almost $15B in savings, bringing Coupa's customers' lifetime savings total to $288B." Read it carefully. That is a platform-wide spend-management figure. It is not savings attributed to AI agents. The release mentions its Navi agents alongside the savings number without claiming the one caused the other, and I am not going to close that gap on the vendor's behalf. Everything else I found in procurement was vendor blogging and consultancy marketing with unfootnoted percentages. IBM named a domain where almost nobody has published a defensible number.

**Customer operations.** This one has the best independent evidence, and it points in two directions at once.

The largest rigorous field study is the NBER working paper on generative AI in customer support, using data from 5,179 agents. It found productivity, measured as issues resolved per hour, rose "by 14% on average, including a 34% improvement for novice and low-skilled workers." That is the shape of the real effect: modest on average, large for the least experienced. It is also the finding most often inflated in marketing decks into a general claim.

Then Klarna, which is the case study everyone cites and almost nobody dates properly. On 27 February 2024 Klarna announced its OpenAI-powered assistant had handled 2.3 million conversations in a month, two-thirds of its customer service chats, "doing the equivalent work of 700 full-time agents," with resolution times falling from 11 minutes to under 2 and an estimated "$40 million USD in profit improvement to Klarna in 2024."

In May 2025 Klarna's CEO conceded that AI customer service, while cheaper, produced lower quality output, and the company restarted hiring human agents. The detail that matters is the timing. That reversal and the workforce-reduction victory lap were six days apart, not years. On 14 May 2025, CNBC reported from Klarna's IPO prospectus that headcount fell "from 5,527 full-time employees as of the end of December 2022 to 3,422 staffers last December."

So Klarna is not a story of a strategy being abandoned. It is a story of a strategy being split: keep the headcount reduction, restore the human escalation path. That is the actual lesson, and it is more useful than either headline. Note also which of those numbers is strongest. The 5,527 to 3,422 figure is in a securities filing. The $40 million is in a press release. When they disagree in tone, believe the filing.

**And then the single best exhibit in the whole file, which is Salesforce against itself.** Salesforce reduced its customer support organisation from about 9,000 people to about 5,000, with its CEO quoted on a podcast saying "I've reduced it from 9,000 heads to about 5,000 because I need less heads." A Salesforce spokesperson confirmed the mechanism was attrition without backfill rather than layoffs, saying "we've seen the number of support cases we handle decline and we no longer need to actively backfill support engineer roles. We've successfully redeployed hundreds of employees into other areas." Salesforce also sells the agent software: Agentforce annual recurring revenue "surpassed half a billion in Q3, up 330% Y/Y," across over 9,500 paid deals and more than 18,500 deals total since launch.

Three months before that headcount comment, Salesforce's own AI researchers published a benchmark of agents on realistic CRM work. Their finding: "leading LLM agents achieve only around 58% single-turn success on CRMArena-Pro, with performance dropping significantly to approximately 35% in multi-turn settings," and the agents "exhibit near-zero inherent confidentiality awareness."

Same company. Same year. Selling the agents, cutting the support staff, and publishing peer-reviewable research showing that leading agents fail roughly two thirds of realistic multi-turn CRM tasks and do not reliably know what they should not disclose. Both statements are primary-sourced. Neither is wrong. They are measuring different things: commercial traction versus task completion.

An independent academic benchmark that simulates an entire company reached a compatible number from the outside, finding "the most competitive agent can complete 30% of tasks autonomously" and that "more difficult long-horizon tasks are still beyond the reach of current systems."

If you want the optimistic end of the range, vendor case studies exist. Intercom publishes a customer story for the prediction market Kalshi showing "80k+ monthly Fin resolutions / 80% automation rate." That is a vendor page about a customer, chosen by the vendor, with no methodology. Weigh it accordingly against a 5,179-agent field study and two benchmarks.

**HR.** IBM's flagship example here is itself. Its AskHR case study reports "a 40% reduction in the HR team's operational costs over the past four years," a "94% containment rate of common questions," and "a 75% reduction in support tickets raised since 2016." It is a strong result and it is IBM measuring IBM, published by IBM, with no external audit. It is also internally slightly untidy: the same page describes "over 2.1 million employee conversations annually" and "more than 11.5 million employee interactions in 2024 alone" without ever defining either unit. Presumably conversations and interactions are different things. The page does not say.

The other HR number worth having is the downside one, and it is court-verified rather than survey-reported. In the EEOC's first AI hiring discrimination settlement, iTutorGroup paid $365,000 after it "programmed their tutor application software to automatically reject female applicants aged 55 or older and male applicants aged 60 or older," rejecting more than 200 qualified applicants. Of the four domains, HR is the one where an automated decision is a regulated act with a named plaintiff class attached. That is not a reason to avoid AI in HR. It is a reason the enterprise version of this work has lawyers in it, and a reason your version should keep a human on the reject decision.

## The place with the best measurement gives the least comfortable answer

Software engineering is where the measurement is genuinely good, and the results are messier than either camp wants.

You have probably seen the METR study that found experienced open-source developers took 19 percent longer to complete issues when allowed to use AI tools, on 16 developers working 246 real issues in repositories averaging 22,000-plus stars and a million-plus lines of code. The perception gap in that study is the memorable part: developers "expected AI to speed them up by 24%, and even after experiencing the slowdown, they still believed AI had sped them up by 20%."

Do not quote that as a present-tense fact. METR itself has stamped its own page: "These results are out of date. We have released results that are current as of early 2026... We believe these historical results no longer reflect the current impact of AI models on open-source developer productivity." The 19 percent figure is a 2025 historical artifact and the people who produced it say so.

Their follow-up, published 24 February 2026, flipped the sign. Returning developers showed an estimated 18 percent speedup and newly recruited developers a 4 percent speedup. Both confidence intervals cross zero, so neither is a result you can lean on. The second experiment also used 10 developers from the original study plus 47 new ones, at $50 an hour instead of $150, a change METR flags as its own confound.

And here is the finding that I think is the most interesting thing in the entire research file, because it is not about model capability at all. METR is redesigning the experiment because it cannot run the old one any more. In their words: "When surveyed, 30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI."

The control group has become unrecruitable. Enough working developers now refuse to do certain tasks unaided that a randomised trial of "with AI versus without AI" biases itself, because participants withhold exactly the tasks where the tool matters most. That is a measurement crisis, and it is also the single strongest evidence of real adoption in the whole file, arrived at by accident. Nobody set out to prove that.

Around that, two large self-selected developer surveys, one run by Stack Overflow and one by Google, measured trust independently and found it moving the opposite way from adoption. Stack Overflow's 2025 survey, with more than 49,000 respondents, found "84% saying they use or plan to use AI tools in their development process, up from 76% in 2024," while "46% of developers said they don't trust the accuracy of the output from AI tools, a significant increase from 31% last year." Forty-five percent said debugging AI-generated code is time-consuming. Google's DORA 2025 report, with around 5,000 respondents, found roughly 90 percent adoption, more than 80 percent reporting productivity gains, a median of two hours a day working with AI, 59 percent reporting a positive influence on code quality, and yet 30 percent trusting it "a little" (23 percent) or "not at all" (7 percent).

The two surveys disagree on the level of distrust, and they should, because they asked different questions with different scales. Treat them as corroborating a direction, not a magnitude. The direction is that adoption and trust are diverging.

Against all of that, IBM says its "entire developer workforce is using Bob, with average productivity gains of 45%," and that AI-enabled internal transformation has delivered "approximately $4.5 billion in productivity savings since the beginning of 2023," with another $1 billion expected in 2026. Alphabet's CEO said in October 2024 that "more than a quarter of all new code at Google is generated by AI, then reviewed and accepted by engineers."

Note the asymmetry in evidence quality. The 45 percent and the $4.5 billion are a company measuring itself with no published methodology, in the same filings where it is selling the capability. METR is an outside lab that measured, published a result it did not want, then publicly retracted the currency of its own finding and explained why its instrument broke. One of those is a marketing number and one is science. They are not the same kind of claim, and the fact that they point in opposite directions should make you slower, not faster.

If you want a failure number that nobody self-reported, use the courts. An academic database of court decisions involving AI-hallucinated legal content lists 1,936 cases worldwide as of 21 August 2026, 1,327 of them in the United States. NPR recorded the same database at 206 cases in July 2025. That is roughly a ninefold rise in thirteen months in documented instances of professionals filing fabricated material in court. In one of those, two attorneys were each fined $3,000 for a filing containing more than two dozen errors including AI-fabricated citations, with the judge writing that the fine was "the least severe sanction adequate to deter and punish defense counsel in this instance."

That curve is the honest counterweight to every adoption statistic on this page.

## What the small-business version of this deal costs

Here is the whole price list, as published on the dates cited. All of it is self-serve.

| What | Price | As of |
|---|---|---|
| ChatGPT Go (US) | $8 per month | 16 January 2026 |
| ChatGPT Plus | $20 per month | 16 August 2026 |
| ChatGPT Pro, lower tier | $100 per month, 5x Plus usage | 20 August 2026 |
| ChatGPT Pro, upper tier | $200 per month, 20x Plus usage | 20 August 2026 |
| ChatGPT Business | $25 per user monthly, $20 annual, 2-seat minimum | 17 August 2026 |
| ChatGPT Business Premium seat (announced, rolling out) | $125 per user monthly, $100 annual, 5x Standard usage | 10 August 2026 |
| ChatGPT Enterprise | Custom, sales only | 21 August 2026 |
| API, GPT-5.6 Sol | $5.00 in / $30.00 out per million tokens | 21 August 2026 |
| API, GPT-5.6 Terra | $2.00 in / $12.00 out per million tokens | 21 August 2026 |
| API, GPT-5.6 Luna | $0.20 in / $1.20 out per million tokens | 21 August 2026 |

Four things about that table that will save you money.

**Prices moved recently and downward.** GPT-5.6 launched on 9 July 2026 at Sol $5/$30, Terra $2.50/$15 and Luna $1/$6. On 30 July 2026, OpenAI cut Luna's price by 80 percent and Terra's by 20 percent. Sol was unchanged. If you built a cost model in mid-July using the launch post, your cheap-tier number is five times too high. Rerun it.

**Batch and Flex are half price, Fast mode is double.** The same Sol model on the Batch or Flex service tier is $2.50 in and $15.00 out per million; in Fast mode it is $10.00 and $60.00. Anything that does not need an answer in the next few seconds belongs on the cheaper tier. Note also that data-residency endpoints carry a 10 percent uplift for models released on or after 5 March 2026.

**Caching is the biggest single lever and it has a gotcha.** Cache reads keep a 90 percent discount, but for GPT-5.6 and later, cache writes are billed at 1.25 times the uncached input rate, with a 30-minute minimum cache life. So caching a prompt you only use once costs you more than not caching it. Cache the stable prefix you reuse all day, not the thing you touch twice.

**Codex is now a seat entitlement, not a separate product.** Standard ChatGPT Business seats include Codex, and standalone usage-based Codex seats stopped being available to new Business workspaces on 24 June 2026. Codex runs across ChatGPT, an editor extension and the terminal on one account. Usage beyond the seat lands in a credit pool: on the Business and Enterprise rate card of 19 August 2026, GPT-5.6 Sol costs 125 credits per million input tokens, 12.50 cached, 750 output, and "a typical Codex task using GPT-5.6 Sol may consume between 5 and 40 credits per task." OpenAI does not publish a dollar-per-credit rate anywhere I could find. The only anchor is its own promotional line offering "$100 worth of workspace credits (2,500 credits)" per Premium seat, which implies four cents a credit. Treat that as inferred from marketing copy, not stated policy.

If you run a nonprofit, note that OpenAI offers up to a 75 percent discount on Business or Enterprise, and ChatGPT for Teachers is free for verified US K-12 educators through June 2027.

Now the honest part, because a piece that ends at "you can rent the same models for $20" is doing the same trick as the vendor decks.

**The integration work does not disappear. It gets much smaller.** IBM's clients are paying for connections into general ledgers, procurement systems, HR systems of record and customer platforms that were bought at different times, run different data models, and in many cases cannot be changed without a change board. That is what the forward-deployed engineer is for. Palantir's framing is exactly right: many capabilities for a single customer.

Your version of that problem is real but it is one or two orders of magnitude smaller. You probably have accounting software, a payment processor, an inbox, a calendar, a CRM or a spreadsheet pretending to be one, and a website. Maybe six systems, all of them current, all of them with documented interfaces, none of them customised in 1998. The work of connecting those to a model is measured in evenings, not in signings.

The other thing the enterprise buys that you cannot replicate is accountability. When an agent misfiles something at a bank, somebody's name is on a statement of work. You do not get that at $20 a seat. What you get instead is the ability to check the output yourself, which at your scale is usually cheaper and always faster.

And you should assume some of it will not work. Both benchmarks above put realistic multi-turn task completion in the 30 to 35 percent range. The field evidence says the gain is largest for the least experienced worker and modest on average. That is a perfectly good deal at $20 a seat. It is a much harder deal to justify at $401.05 an hour.

## What I would actually do, in order

1. **Pick one process with a countable unit before you buy anything.** Invoices coded per hour. Quotes drafted per day. Support conversations resolved without escalation. If you cannot count it now, you will not be able to prove anything later, and every number in this article will be untestable in your business.
2. **Measure the baseline for two weeks with no AI in the loop.** This is the step everyone skips and it is the reason so much of the published evidence is worthless. METR's whole experimental design collapsed because people would not work unaided any more. Get your baseline while you still can.
3. **Buy two ChatGPT Business seats billed annually, at $20 per user per month.** Two is the minimum. That is $480 for the year, less than 72 minutes of an IBM Consultant V at the 2026 GSA ceiling of $401.05 an hour. Codex is included in the seat.
4. **Put the cheap model on the boring half.** Route bulk classification, extraction and summarisation to GPT-5.6 Luna at $0.20 in and $1.20 out per million tokens, and reserve Sol at $5 and $30 for the work that actually needs it. Run anything non-interactive on Batch or Flex at half price.
5. **Cache only the stable prefix.** Cache reads are 90 percent off, but cache writes cost 1.25 times uncached input with a 30-minute minimum life, so caching one-shot prompts loses money.
6. **Keep a human on any decision that creates a legal duty.** Hiring rejections, credit decisions, anything filed with a court or a regulator. The EEOC's first AI hiring settlement was $365,000 for an automatic age filter, and the running count of court decisions involving AI-fabricated content went from 206 in July 2025 to 1,936 in August 2026.
7. **Re-measure the same countable unit after 60 days and compare it to the baseline, not to your feelings.** In METR's original trial, developers who had been measurably slowed still believed they had been sped up by 20 percent. You will not be immune to that and neither am I.
8. **Only then consider paying anyone.** If after all that you still have a genuine integration problem, you now know exactly what it is, which is the single best position from which to buy help.

If you are already paying a retainer for work in this territory and want the full argument about what that money buys, [The $20 Dollar Agency](https://the20dollaragency.com/) is the long version. This post is the short one.

## The point

IBM is doing something entirely rational. Its clients have decades of accumulated systems, real regulatory exposure, and a genuine need for somebody to be accountable when an agent gets it wrong. Certified consultants and forward-deployed engineers are the correct answer to that problem, and $401.05 an hour is not an outrageous price for it given an 11.7 percent segment margin.

It is just not your problem. Your systems are new, your data is in five places instead of five hundred, and the models in that partnership are on a self-serve page with a credit card form. The fragmentation is the product. You do not have any.

## Fact-check notes and sources

- **The partnership itself**: announced 13 August 2026, naming finance, procurement, customer operations and HR, embedding GPT-5.6, Codex and ChatGPT Work into IBM Consulting Advantage, and launching a dedicated OpenAI Practice. [IBM newsroom, 13 August 2026](https://newsroom.ibm.com/2026-08-13-ibm-partners-with-openai-to-accelerate-secure-ai-deployment-for-enterprises-across-core-operations). The "Elite tier" claim comes from IBM's release; see the tier note below.
- **IBM segment results for the quarter ended June 2026** (Consulting $5.3 billion flat, Software $7.8 billion up 5 percent, Infrastructure $3.8 billion down 7 percent, total $17.2 billion up 1 percent, full-year guidance of four-to-five percent constant currency growth): [IBM Releases Second-Quarter Results, 22 July 2026](https://www.prnewswire.com/news-releases/ibm-releases-second-quarter-results-302832559.html). IBM pre-announced preliminary results on [14 July 2026](https://newsroom.ibm.com/2026-07-14-Arvind-Krishnas-Letter-to-IBM-Investors), eight days early, with a note that final figures could differ.
- **IBM Consulting full-year 2025 figures** (revenue $21,055 million, segment profit $2,464 million at an 11.7 percent margin, signings $21,757 million down 13.3 percent, book-to-bill 1.03, backlog $31.9 billion, generative AI book of business above $12.5 billion inception to date, total company revenue $67.5 billion, free cash flow $14.7 billion, approximately $4.5 billion of productivity savings since the beginning of 2023): [IBM 2025 Annual Report, filed as Exhibit 13 to the Form 10-K, 24 February 2026](https://www.sec.gov/Archives/edgar/data/51143/000005114326000027/ibmars2025.pdf).
- **The definition of the generative AI book of business** (inception-to-date Software transactional revenue plus new SaaS annual contract value plus Consulting signings, roughly one-fifth Software and four-fifths Consulting): [IBM Form 8-K Exhibit 99.2, 22 July 2026](https://www.sec.gov/Archives/edgar/data/51143/000005114326000077/ibm-20260722xex992.htm). The switch to percentages ("about 50% of our signings", "over 30% of our backlog") and the 6 percent signings growth: [IBM 2Q26 prepared remarks](https://www-api.ibm.com/adobe/assets/urn:aaid:aem:023c55fc-8c53-428f-a5cf-42cb35bb78db/original/as/ibm-2q-26-earnings-prepared-remarks.pdf). I read the full 1Q26 and 2Q26 prepared remarks; the phrase "book of business" appears in neither. The $9.5 billion figure that circulates online is a real but earlier IBM number and should not be used undated.
- **IBM Bob and the internal productivity claims** (entire developer workforce, 45 percent average productivity gain, $1 billion additional savings expected in 2026): [IBM 1Q26 prepared remarks, 22 April 2026](https://www-api.ibm.com/adobe/assets/urn:aaid:aem:608df6c6-4d5b-4169-85b5-fda213156bf3/original/as/ibm-1q-26-earnings-prepared-remarks.pdf). These are IBM measuring IBM with no published methodology. The widely circulated "$3.5 billion" version of the savings figure is not what IBM says; IBM's own [Think page](https://www.ibm.com/think/insights/enterprise-transformation-extreme-productivity-ai) states USD 4.5 billion.
- **IBM headcount**: not disclosed. The FY2025 10-K defers human capital disclosure to the annual report, and that section contains no employee count for IBM or for Consulting. The only staffing figure anywhere in it is that over 200,000 employees participated in an engagement survey in 2025, which is a participation count, not a total. Figures like "160,000 IBM consultants" circulate widely; I found no IBM-published source and have not used one.
- **IBM's published hourly rates**: GSA Multiple Award Schedule contract [GS-35F-110DA](https://www.gsaelibrary.gsa.gov/ElibMain/contractorInfo.do?contractNumber=GS-35F-110DA&contractorName=INTERNATIONAL+BUSINESS+MACHINES+CORPORATION&executeQuery=YES), current option period ending 20 December 2030. IT-services labour rates ($401.05 Consultant V, $309.62 Consultant III, $248.34 Consultant I, $386.75 Architect V for 2026) are in [Appendix C](https://www-api.ibm.com/adobe/assets/urn:aaid:aem:db7fe7a6-3169-475e-896a-bfaa6bc706e7/original/as/appendix-c_it-services-technical-and-consulting-labor-rates-and-descriptions.xlsx); the management-consulting card with Consultant I at $109.46 is [Appendix C.5](https://www-api.ibm.com/adobe/assets/urn:aaid:aem:e1ae4af1-0b5c-44be-ae08-b141ab1f513f/original/as/appendix-c-5-professional-services-labor-rates-and-descriptions.xlsx). These are ceiling prices including the Industrial Funding Fee, so actual orders are at or below them. The $3.1 million figure is my own arithmetic from the published Consultant III rate, not an IBM statement. IBM does not appear in GSA's contract-awarded labour rate database at all; its rates live only in these schedule documents.
- **Accenture comparison figures** (Q3 FY2026 revenue $18.72 billion, new bookings $19.32 billion down 2 percent, Consulting revenue $9.3 billion, approximately 799,000 people): [Accenture Q3 FY2026 results, 18 June 2026](https://newsroom.accenture.com/content/3qfy26-earnings/accenture-reports-third-quarter-fiscal-2026-results.pdf) and the [call transcript](https://investor.accenture.com/~/media/Files/A/accenture-v4/investors/earnings-reports/2026/accenture-third-quarter-fiscal-2026-conference-call-transcript.pdf). Its last published generative AI figures are FY2025: $5.9 billion of bookings against $80.6 billion total, per the [FY2025 results release, 25 September 2025](https://newsroom.accenture.com/content/4q-full-fy25-earnings/accenture-reports-fourth-quarter-and-full-year-fiscal-2025-results.pdf), and $2.7 billion of revenue against $69.7 billion total, plus the 77,000 AI and data professionals and 6,000-plus projects, per the [FY2025 call transcript](https://investor.accenture.com/~/media/Files/A/accenture-v4/investors/earnings-reports/2025/accenture-fourth-quarter-fiscal-2025-conference-call-transcript-updated-10-8-2025.pdf). The June 2026 release and transcript carry no advanced-AI bookings or revenue figure.
- **OpenAI Partner Network** (three tiers, $150 million investment, 300,000 certified consultants targeted by end of 2026, specializations, Forward Deployed Experts pilot): [Introducing the OpenAI Partner Network, 14 June 2026](https://openai.com/index/introducing-openai-partner-network/). On tier labels: OpenAI's [partner entry for IBM](https://openai.com/business/partners/ibm/) shows countries, industry and program partners, with no tier badge, and I found no published tier thresholds or partner tier assignments anywhere on OpenAI's site. The Elite claim in this article is IBM's own.
- **Forward deployed engineering**: OpenAI's [Deployment Company launch, 11 May 2026](https://openai.com/index/openai-launches-the-deployment-company/) states more than $4 billion of initial investment, 19 partner firms led by TPG, and approximately 150 forward deployed engineers and deployment specialists arriving via the Tomoro acquisition. A third-party blog circulates a "$10 billion" figure and a guaranteed annual return for that vehicle; neither appears on any OpenAI page and I have not used them. The origin of the role is Palantir's own [2020 description of the Forward Deployed Software Engineer](https://blog.palantir.com/a-day-in-the-life-of-a-palantir-forward-deployed-software-engineer-45ef2de257b1).
- **ChatGPT Work is an agent, not a plan**: [OpenAI, 9 July 2026](https://openai.com/index/chatgpt-for-your-most-ambitious-work/). The commercial plans remain ChatGPT Business and ChatGPT Enterprise.
- **All consumer and business prices** come from OpenAI Help Center articles and the [business pricing page](https://openai.com/business/chatgpt-pricing/), not from the main pricing page, which loads its prices in the browser rather than printing them on the page, so there is no figure on it to quote. Sources: [ChatGPT Business, 17 August 2026](https://help.openai.com/en/articles/8792828-what-is-chatgpt-business) for the $25/$20 seat, the two-seat minimum, the 29 August 2025 rename from ChatGPT Team, and the 24 June 2026 closure of standalone Codex seats to new workspaces; [Premium seats, 10 August 2026](https://openai.com/index/premium-seats-chatgpt-business/) for $125/$100 and the "$100 worth of workspace credits (2,500 credits)" line that implies four cents a credit; [ChatGPT Plus](https://help.openai.com/en/articles/6950777-what-is-chatgpt-plus) for $20; [Pro tiers, 20 August 2026](https://help.openai.com/en/articles/9793128-about-chatgpt-pro-tiers) for $100 and $200 and the absence of annual billing on Go, Plus and Pro; [ChatGPT Go, 16 January 2026](https://openai.com/index/introducing-chatgpt-go/) for $8 in the US; and the [ChatGPT pricing page](https://openai.com/chatgpt/pricing/) for the nonprofit discount and the teacher plan.
- **API token prices and service tiers**, observed 21 August 2026: [OpenAI API pricing docs](https://platform.openai.com/docs/pricing). The launch prices and the caching rules are in the [GPT-5.6 announcement of 9 July 2026](https://openai.com/index/gpt-5-6/), which also carries the 30 July 2026 update note recording the 80 percent Luna cut and 20 percent Terra cut. Credit rates and the 5 to 40 credits per typical Codex task are from the [ChatGPT rate card, 19 August 2026](https://help.openai.com/en/articles/20001106-codex-rate-card).
- **Finance and procurement evidence**: BlackLine's "up to a 92% reduction" is a vendor claim about self-selected early adopters, from its [general availability release of 27 July 2026](https://blacklinesystemsinc.gcs-web.com/news-releases/news-release-details/blackline-advances-governed-ai-finance-general-availability/). Coupa's $425 billion of managed spend and almost $15 billion of savings come from its [quarterly release of 15 December 2025](https://www.prnewswire.com/news-releases/coupas-ai-driven-procurement-saves-customers-almost-15b-in-q3-fy26-302642106.html) and are platform-wide spend-management figures, not savings attributed to AI agents. I looked for a defensible primary procurement number beyond this and did not find one.
- **Customer operations**: the 14 percent average and 34 percent novice gains are from [NBER Working Paper 31161, Generative AI at Work](https://www.nber.org/papers/w31161), a field study of 5,179 support agents, later published in the Quarterly Journal of Economics. Klarna's original claims are in its [27 February 2024 release](https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/); the quality reversal and rehiring were reported on [9 May 2025](https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396); the IPO prospectus headcount figures were reported by [CNBC on 14 May 2025](https://www.cnbc.com/2025/05/14/klarna-ceo-says-ai-helped-company-shrink-workforce-by-40percent.html). Salesforce's support headcount comment and its spokesperson's attrition explanation were reported by [The Register, 2 September 2025](https://www.theregister.com/2025/09/02/salesforce_4000_jobs_ai/); Agentforce commercial figures are from [Salesforce's Q3 FY26 release, 3 December 2025](https://www.salesforce.com/news/press-releases/2025/12/03/fy26-q3-earnings/); the agent benchmark from its own researchers is [CRMArena-Pro, arXiv:2505.18878](https://arxiv.org/abs/2505.18878). The 30 percent autonomous completion figure is from [TheAgentCompany, arXiv:2412.14161](https://arxiv.org/abs/2412.14161). The Kalshi figures are a vendor customer story on [Intercom's customers page](https://www.intercom.com/customers) with no stated methodology.
- **HR**: IBM's AskHR results, including the internal mismatch between "conversations annually" and "interactions in 2024," are on [IBM's own case-study page](https://www.ibm.com/case-studies/ibm-askhr). The EEOC settlement is [iTutorGroup, 11 September 2023](https://www.eeoc.gov/newsroom/itutorgroup-pay-365000-settle-eeoc-discriminatory-hiring-suit).
- **Software engineering measurement**: the 19 percent slowdown, the 16-developer sample and the perception gap are in [METR's July 2025 study page](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/), which now carries METR's own notice that the results are out of date. The follow-up estimates (18 percent and 4 percent speedups, both with confidence intervals crossing zero), the 10 returning plus 47 new developers, the pay change from $150 to $50 an hour, and the 30 to 50 percent of developers withholding tasks are in [METR's 24 February 2026 design update](https://metr.org/blog/2026-02-24-uplift-update/). METR's [May 2026 self-report survey](https://metr.org/blog/2026-05-11-ai-usage-survey/) found a median self-reported 1.4x to 2x value change and a 3x speed change, which is a self-report, not a measurement. Trust figures: [Stack Overflow 2025 Developer Survey](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/), more than 49,000 respondents, and [Google's DORA 2025 report](https://blog.google/innovation-and-ai/technology/developers-tools/dora-report-2025/), around 5,000 respondents. The two use different question wordings and are not directly comparable on magnitude. Alphabet's AI-generated code share is from [Sundar Pichai's Q3 2024 earnings remarks](https://blog.google/inside-google/message-ceo/alphabet-earnings-q3-2024/). A comparable Microsoft figure circulates widely but I could not source it to any Microsoft-owned page, so it is not here.
- **Court-verified AI failures**: the [AI Hallucination Cases database](https://www.damiencharlotin.com/hallucinations/) listed 1,936 cases worldwide and 1,327 in the United States when I checked it on 21 August 2026. [NPR recorded the same database at 206 cases on 10 July 2025](https://www.npr.org/2025/07/10/nx-s1-5463512/ai-courts-lawyers-mypillow-fines) and quotes the $3,000-per-attorney sanction order directly. The database's compiler notes it is not exhaustive, so the real figure is higher.
- **Two claims I deliberately left out**: a widely quoted forecast that a large share of agentic AI projects will be cancelled within a couple of years reached me only through secondary reporting, and a widely quoted study claiming that almost all enterprise generative AI pilots produce no measurable return rests on a few dozen interviews and a few hundred survey responses that I could not read in the original. Neither is load-bearing here, so neither is cited as fact.

## Related reading

- [The $50-a-month AI stack for a small business](/blog/blog-50-month-ai-stack-smb/): the concrete build that sits underneath the price table above.
- [Running AI agents without a surprise bill](/blog/blog-ai-agent-cost-controls-smb/): the cost controls to put in place before you turn anything on.
- [Hire, upskill, or outsource your AI work](/blog/blog-smb-hire-upskill-outsource-ai/): the same buy-versus-build decision at your scale rather than IBM's.
- [A study like this quotes at $15,000](/blog/market-study-cost-human-versus-ai/): what the consulting price actually covers when you take it apart.
- [Wall Street is betting on the wrong AI layer](/blog/blog-boring-ai-layer-decides-smb-margin/): why the unglamorous integration layer decides your margin, not the model.

*This post is informational and is not legal, financial or investment advice. All figures are as published on the dates cited and change frequently, particularly model prices. Company statements quoted here are those companies' own claims unless a source is identified as independent. No affiliation with IBM, OpenAI or any other vendor mentioned is implied, and all mentions are nominative fair use.*


---

Canonical HTML: https://jwatte.com/blog/ibm-openai-what-large-firms-are-buying/
RSS: https://jwatte.com/feed.xml
JSON Feed: https://jwatte.com/feed.json
Hero image: https://jwatte.com/images/ibm-openai-what-large-firms-are-buying.webp
