Is AI a Bubble? · Dated investigation
Cheap extraction is not the same as cheap accepted work
Efficiency can make a useful service cheaper. Who pays for the checking, integration and continuing work—and who keeps what is left?
The buyer pays for usable content, not just generated pages. A released model operated on rented computing capacity is a credible alternative for some bulk work. The economic advantage survives only if the extra checking, service and continuing costs fit inside its saving. The numerical tests below solve those allowances; they do not establish that a particular customer or provider achieved them.
1. The answer: efficiency can pay, but the recipient and remaining work decide how much
There is a stronger, concrete case for productive AI efficiency than for either universal price collapse or universally high supplier returns. A specialized open-weight system can convert documents cheaply enough to be a credible substitute for a paid extraction component. That does not make the complete workflow free, establish that every customer should self-host, or reveal a frontier lab's owner cash. The relevant comparison is the cost of content accepted for a particular use after checking and integration—not the price of generating another candidate page.
The principal evidence is Ai2's documented olmOCR service and development process, compared with Mistral's separately priced OCR offerings. A reported timed batch supports a rental-cost calibration of about $183 per million submitted pages at an October-inspected H100 offer. A current OCR3 batch offer costs $1,000 per million pages before customer-specific adjustments. These are different cost boundaries: the first is a modeled machine-rental component; the second is an offered service price. The approximately $817 difference is the maximum extra per-million-page burden that self-hosting could absorb under equal locally established adequacy and other equal costs. It is not an observed saving. At an assumed $40 per hour of review resources, it corresponds to only 0.0735 additional seconds per input page, or one additional thirty-second intervention per approximately 408 pages.[1][5][6][8]
The supplier calculation is equally consequential. In a stipulated year with 80 million input pages at $0.001 each, all collected, continuous rental of one active H100 at the inspected $2.99/hour offer and an incomplete $391.90 development reference leave $53,416 for all remaining costs and any surplus. That volume fits the historical throughput only under the explicit 75% productive-time ceiling; it is not measured demand or uptime. Adding a passive hot spare with no extra productive capacity and an assumed average 0.05 seconds of extra operator-paid review per input at $40/hour makes that residual −$17,221. The difference is a concrete change in required service resources, not a claim that the model's observed accuracy deteriorated. Full ongoing development, support, security, taxes and other costs still have to fit inside the positive residual; they are not presumed absent.
This changes the whole-pool interpretation in three ways. First, efficiency can sustain useful services and positive contribution even where a prior generation of financing or ownership expectations disappoints. Second, cheaper ingredients may principally benefit the buyer or the application owning the workflow; they need not preserve the old model provider's receipt per task. Third, the ability to expand volume is a finite capacity and commercial question. It cannot be assumed to repair every price cut. A metered coding service supplies a contrasting operative response: GitHub changed the billing unit to reflect token-consuming work, rather than promising the same effective entitlement for every increasingly long agent session.[10][11]
The strongest favorable account is therefore selective and economically coherent: adequate specialized service, a large repeatable workload, shared overhead, bounded intervention, and a pricing arrangement that finances continuing quality. The strongest adverse account is not that AI stops working. It is that competition compresses the money left per accepted task faster than utilization, automation and cost control improve, while reliability and development costs remain. Both are compatible with continuing customer value. The evidence does not yet identify the distribution of these outcomes across the market.
2. What is being bought—and why the earlier buyer cases are not the numerical anchor
2.1 The accepted unit
Consider a document team building a searchable technical archive or a training-data corpus. It needs text in the right order, tables with the relevant cell relationships preserved, usable equations where required, and links back to the source. A file that completes an inference call but silently moves a value to the wrong table row has not delivered that work. Nor has an apparently correct page delivered the full service before its output is ingested, checked according to the application's risk, and made retrievable.
Here, accepted work means content that passes the buyer's specified task-level checks after any necessary repair or fallback. The study does not invent a universal threshold that makes every legal, financial, archival or scientific use adequate. A corpus builder may legitimately discard an unusable page. A team processing a fixed set of required invoices cannot replace a failed invoice with an unrelated easy page. The latter must pay for repair, another service or manual completion. These are different denominators and different costs.
For an eligible corpus in which rejection is permitted, define q as retained acceptable pages divided by submitted pages. To obtain A acceptable pages, the arithmetic input requirement is A/q, provided the retained collection still meets the buyer's content requirements. Every failed input still consumes resources. For a fixed document obligation, track each required document's complete processing and fallback cost instead; do not use A/q to disguise missing documents. The companion includes a clearly assumed 98% denominator illustration, not an empirical acceptance estimate, and the principal operator budget does not use it.
2.2 The parties and the allocation of benefit
The payer is the organization purchasing conversion or funding an internal processing team. The user is the data, document or review team; it may not control the budget. The application operator combines extraction, storage, checking and integration. The model developer supplies weights or a managed endpoint; the capacity supplier sells compute time. A buyer's smaller external bill, a reviewer's reduced workload and a provider's cash surplus are separate gains. Adding them as if they were three independent revenues would count the same improvement more than once.
Self-hosting here means operating released weights on rented capacity. It does not mean owning the GPU. The operator takes responsibility for the surrounding workflow while the cloud supplier retains its hardware and facility economics. A managed OCR API can bundle part of that operational work into its price. Neither option automatically includes the buyer's content-specific acceptance decisions, compliance obligations or final integration.
Accepted-work economics · evidence and assumptions distinguished
Start with the work someone can actually use
On a narrow screen, scroll within the figure to inspect the labels.
2.3 Why this is a different workload from the earlier buyer cases
the buyer-budget investigation's council contract and marketing-production account establish actual buying and budget substitution at their own boundaries. They do not supply a matched set of input files, competing model outputs, billed serving hours, acceptance rates or incremental review times. A recording hour is not a PDF page; a final campaign asset is not extracted text. Transferring the council's reported assessment increase into an OCR quality parameter would be a fabricated match.[15]
The document-conversion alternative is stronger for this chapter because it offers a public timed run, a development account, downstream evidence, released model information and priced alternatives. It is weaker on commercial final demand: Ai2's large deployment is its own model-training input, not evidence that independent outside customers paid the scenario's page revenue. The hypothetical commercial volume is kept separate from the observed operating capability. The comparison answers the full-cost efficiency question without claiming that the earlier buyers use olmOCR or any specified campus.
Interactive coding provides a second, narrower comparison. Its useful output is an integrated, tested change—not an autonomous session or a token. The evidence there concerns an actual change in the seller's billing and cost-risk allocation, not a newly measured coding productivity effect. It tests portability without turning the chapter into a model leaderboard.
3. What the quality evidence supports
3.1 A credible alternative, not certified equivalence for every buyer
The olmOCR work uses tests of text, reading order and structural relationships. Its reported score is a proportion of tests passed, averaged across categories; it is not the fraction of whole pages a particular customer would accept. The suite evolves with discovered failures. That is useful for development, but it is not an unchanged independent holdout across every release.[1]
There is also a more economically relevant historical check than a standalone extraction score. The earlier paper reports a same-corpus continued-pretraining comparison in which the average downstream score moved from 53.9 to 55.2 with the olmOCR-derived corpus rather than the Grobid-derived corpus. The calculated improvement is 1.3 score points, not a worker-output percentage or a cash return; individual components did not uniformly improve. Its relevance is narrower and positive: the alternative extraction was useful in that tested downstream setting, rather than simply cheaper at producing any text.[2]
The model card identifies a full toolkit, including rotation, rendering, metadata handling and retries. Calling hosted weights without the surrounding pipeline is not necessarily the same product. The research therefore uses a timed pipeline batch as its compute anchor, not a peak token rate converted through an invented average document length.[4]
For the present comparison, open-model adequacy is established only as a credible option for evaluation, supported by its scoped tests and use. Current Mistral versions are not the historical API version in the older benchmark. No current, buyer-specific head-to-head acceptance study was obtained. The conditional economic thresholds below tell a buyer or operator how good the match must be financially; they do not certify that the match has occurred.
3.2 Paying more can purchase a different accepted result
Mistral's OCR3 model card, dated December 18, 2025 and inspected at this cutoff, lists $2 per 1,000 ordinary pages, with OCR3 still available for existing integrations. Its newer OCR4 offering lists $4 per 1,000 ordinary pages and adds structured output features such as source boxes and confidence information. Document AI or annotation layers are not automatically included in an ordinary-extraction price.[6][7]
The Batch API provides a 50% price reduction in an asynchronous mode. Thus the selected batch references are $1 per 1,000 for OCR3 and $2 per 1,000 for OCR4. These are offers, not average realized receipts, and a batch job is not a promise of synchronous low-latency interaction.[8]
A buyer that needs better provenance or fewer consequential extraction mistakes may rationally buy the more expensive version. Under equal volume and other costs, the OCR4-versus-OCR3 batch premium is $0.001 per input page. At the report's assumed $40/hour resource value, avoiding 0.09 seconds of review per page would offset it. That is a required reduction, not an observed OCR4 effect. A binding security, accessibility or accuracy requirement may also justify the choice without a payroll saving; the requirement still needs evidence rather than a favorable label.
This service-version comparison prevents an erroneous market story. A newer AI service can be priced higher because it does different work. Conversely, leaving the older service available preserves a lower-priced route for users who do not need the additional features. Neither observation establishes a uniform quality-adjusted price trend. The provider itself qualifies its benchmark comparisons and recommends document-specific evaluation.[7]
3.3 What has to remain matched
The numerical core is offline, bulk document processing, with enough queued work to use a single GPU efficiently. It does not cover urgent first-token latency, a guaranteed response deadline, very long interactive context, arbitrary multilingual handwriting, or unmeasured concurrent customer bursts. Document rendering, rotations, repeated outputs, output length and hardware precision can all change throughput. Exact sustained live concurrency and the cost of the buyer's final checks are not observed.
The measured elapsed time already incorporates the processing done in that experiment. There is no additional hypothetical caching discount or second retry speed-up applied to it. Duplicate-document caching can reduce work only where content, permissions and customer billing permit reuse; the present budget assumes no such saving. Cross-document deduplication can also reduce billable volume, so it cannot automatically be claimed as both the same sales and less work.
For a deployed comparison, the relevant fields are a pinned model and pipeline, rendering resolution, input and output length distributions, timeout and retry rules, batch size/concurrency, acceptance tests, review and fallback, and the retention/security requirement. Public evidence covers some of these, not all. The missing fields are why the report supplies a conditional full-cost boundary rather than an audited cost per approved customer document.
4. Reconstructing the serving cost instead of copying the printed headline
4.1 An internal denominator conflict
The July 2026 ACL paper reports 10,000 pages in 36 minutes 47 seconds on one H100 using FP8, with a historical rental assumption of $2.69/hour. Yet Table 1 and §5 label approximately $165 as the cost per 10,000 pages. The appendix arithmetic instead gives:
2,207 seconds / 3,600 × $2.69 = $1.649119 for 10,000 pages
$1.649119 × 100 = $164.911944 per million pages.[1]
The same inconsistency appears in the Mistral comparison: the appendix's historical $1 per 1,000 pages implies $10 for 10,000 pages, whereas the comparative table lists $1,000 under its stated denominator. Both printed tables and the runtime passage were checked visually. Ai2's October 2025 release description separately says fewer than $2 for 10,000 pages, consistent with the timed-rental calculation's scale.[3]
Research inference: the printed denominator is likely mislabeled, but no author clarification or formal correction was obtained. The ledger preserves the conflicting amounts as source_conflict; the model does not quietly reinterpret them as observations. It uses the explicitly reported runtime and rate. This is not a newly observed hundredfold change in cost, a finding about invoices, or a correction to earlier Shaduf financial research.
The reconstruction matters to the economic test. Taking $165/10,000 literally would make the machine-rental component exceed the current batch service offer by a large multiple. Reconstructing the stated experiment instead leaves a substantial—but finite—allowance for everything beyond compute. The inference rests on the source's own timing and prices, not a preferred conclusion about open models.
4.2 A dated quote applied to a dated experiment
RunPod's inspected Pods table offers an H100 SXM 80GB at $2.99/hour. Applying that different price to the historical 2,207-second batch gives:
c = (2,207 / 3,600 × $2.99) / 10,000 = $0.000183303611 per submitted page.
That is $183.303611 per million input pages. It is a hybrid rental calibration, not a rerun on an October workload, a complete GPU ownership cost or the provider's achieved receipt. The offer's included host resources, deployment availability and separate storage services do not establish that the original test environment can be reproduced at that exact all-in bill.[5]
The historic launch test supplies a consistency check, not a matched performance trend: its different 1,288-page run reconstructs to about $175.49 per million using the then-cited L40S rate, and $178.10 using the then-cited H100 rate. The accompanying general-API test reconstructs to $16.0693 for that 1,288-page sample at the paper's historical token prices. Those token counts and prices are not transferred to a newer model or called current offers.[2]
Accepted-work economics · evidence and assumptions distinguished
Read the denominator before using the price
On a narrow screen, scroll within the figure to inspect the labels.
4.3 Finite productive capacity
The newer batch implies 16,311.735 submitted pages per elapsed hour under its test conditions. Straight multiplication over 8,760 hours gives approximately 142.89 million pages. That is an arithmetical ceiling, not continuously observed operation. The model imposes a 75% productive-time ceiling, leaving a chosen allowance for interruptions and peaks:
N_max = 10,000 / (2,207 / 3,600) × 8,760 × 0.75 = 107.168 million input pages/year.
The 75% is a Research assumption, not measured GPU utilization, a customer service-level agreement or a source-derived downtime rate. Real document mix, host bottlenecks, load variation or maintenance may make capacity lower. A hot spare in the reliability stress adds paid availability capacity but is deliberately not counted as simultaneously productive throughput. Workloads exceeding the stated active capacity are rejected by the calculation rather than served for free.
5. The buyer's threshold: how much extra work can the cheap component afford?
At the OCR3 batch reference, one million submitted pages cost $1,000 in API charges. The active-time rental component for the open pipeline is $183.30. Under equal accepted output and common costs, the maximum extra burden that can be paid before self-hosting loses that component-price advantage is:
Extra setup + incremental review/rework + extra hosting + migration + incremental development ≤ $816.696389 per million inputs.
The comparison includes only differences in the two complete workflows. Review needed under both options cancels only when it is genuinely equal. A managed API that requires more checking can be worse despite outsourcing operations; an open pipeline with more checking can be worse despite cheaper generation. Neither direction is assumed.
At an assumed fully loaded review-resource value of $40/hour:
$0.000816696389 / ($40 / 3,600) = 0.073502675 extra seconds per input page.
This is approximately one additional thirty-second intervention every 408 pages. It is not a claim that a human can review a page in 0.0735 seconds. It states how little average extra work would consume the entire component-price difference. Five hundred difficult pages can absorb more money than a very large number of nearly free easy pages save. The actual tail of difficult cases, not a generic average benchmark score, determines whether this happens.
Fixed implementation matters as well. A one-time migration or continuing support burden should be divided over the relevant committed, adequately processed workload—not whatever global volume makes the spread attractive. The model's active-time comparison is generous to self-hosting because it does not yet charge a dedicated idle GPU. Section 6 supplies that stricter operator boundary. A low-volume organization with an existing shared team and flexible batch scheduling faces a different incremental choice from a new vendor supporting customer-specific contracts around the clock.
For permitted corpus rejection, the generic accounting relation is:
Cost per accepted page = (input-unit cost + input-unit review/rework cost) / q + other attributable cost / A.
The 98% example in the code would require 1.020408 million inputs for a million acceptable retained pages, and $1,020.41 of OCR3 batch charges before other costs. It is included only to make failed inputs nonfree. It is not a measured acceptance rate or a shortcut for a fixed obligation in which every original document must be completed. When rejection is not permitted, price the actual repair/fallback path for the required set.
Accepted-work economics · evidence and assumptions distinguished
How much extra work can the cheaper component afford?
On a narrow screen, scroll within the figure to inspect the labels.
What this establishes positively: the public cost evidence identifies a small, inspectable resource allowance. A buyer does not need a lab's audited corporate margin to test whether a proposed migration can fit inside it. But without matched acceptance and review observations, the analysis cannot certify which vendor is cheaper for that buyer. The next measurement should be accepted output and the extra work needed to obtain it—not another posted token price.
6. A supplier can cover serving while still failing to fund the business
6.1 Put the unmeasured costs inside the question, not outside the conclusion
An operator may sell extraction at a per-page price while paying both usage-dependent and standing costs. The appropriate question is how much money remains to finance everything not already charged. Define:
H = k × p × N − rental cash − priced development reference − explicitly charged review.
Here N is submitted, billable page volume; p is the stipulated price; k is its net collected fraction; and H is the remaining-cost budget. Actual other costs must be subtracted from H before a surplus is established. If H is positive, it is the maximum room for those costs plus any distributable residual—not an audited profit, a gross-margin disclosure or a forecast of owner cash. If negative, the named modeled uses already exceed the modeled receipts; it does not establish an actual missed payment.
The remaining costs include continuing model work and data, unpriced experiments and evaluation, product engineering, customer acquisition and support, operational security and monitoring, additional storage and host services, migration, corporate overhead, taxes, working-capital finance, and renewal spending not already included in the rent. They must also include economically relevant compensation rather than counting an employee's work as free because payment or accounting treatment differs. An operator with shared teams can have a small incremental allocation; a standalone provider cannot borrow that favorable allocation without evidence.
The model uses a collected price of $0.001 per input page, equal to the inspected OCR3 batch offer, solely as a reference competitive price. No evidence establishes that the hypothetical operator has customers at that price, achieves the managed API's service level, or collects every billed dollar. The base k=1 is a favorable assumption. Discounting net receipts by 10%, without changing the workload, lowers the main residual by $8,000.
6.2 Pay for continuing development without pretending its total is known
The newer paper reports a 22-GPU-hour supervised fine-tune on one B200, and 2,186 synthetic pages at an estimated $0.12 each. It separately describes an eight-H100 reinforcement-learning configuration and multiple seeds without the complete elapsed program cost. Teacher labeling, base-model development, human filtering and the full engineering program are not priced by those two observations.[1]
For a transparent partial charge, value the single fine-tune at the inspected $5.89/B200-hour offer and retain the reported synthetic-data estimate:
22 × $5.89 + 2,186 × $0.12 = $129.58 + $262.32 = $391.90.[5]
The scenario pays that subset once per modeled year. Annual repetition is an analytical convention, not observed maintenance cadence. The quoted rental configuration is not independently shown to reproduce the reported training environment, and this is not the developer's invoice. Most importantly, $391.90 is not the price of creating or maintaining the released model. All remaining competitive development must fit inside H. The deliberately small priced subset prevents the exercise from claiming a zero cost while exposing how little of the full program is actually measured.
The earlier project's disclosure is a useful warning against a common shortcut: 365 training-node hours for the experiments versus 16 for one final run, a ratio of 22.8125. The smaller amount is part of the larger one, not an extra charge to add. That dated program is not transplanted into the newer annual budget, but it demonstrates why multiplying a final-run duration by an hourly price need not price the effort required to obtain the result.[2]
A user of released weights may legitimately choose not to train a new model. That avoids its own training bill; it does not establish that competitive updates cost nobody anything. Continued external releases, paid maintenance, in-house adaptation and keeping an adequate older version are different strategies. Each can work, and each carries different obligations. Open availability shifts who must finance development; it does not erase the development function. The base model's historical cost is sunk for a present usage decision, while future adaptation or replacement may remain necessary.[4]
6.3 The rental account and its capacity boundary
Continuous rental of one active H100 at $2.99/hour costs:
8,760 × $2.99 = $26,192.40/year.
At 80 million pages, the stipulated receipts are $80,000. They use about 56% of the raw annual timed capacity and fit beneath the chosen 75% ceiling. The base annual balance is:
$80,000 − $26,192.40 − $391.90 = $53,415.70.
That amount could support remaining incremental costs in a lean, shared operation. It cannot be called a sustainable business return until those costs are established. The calculation does not require the operator to buy a second GPU merely because demand is higher than average, but it also does not promise that a single machine supplies uninterrupted service.
| Conditional annual operator case | Receipts | GPU rental | Priced development subset | Additional review | Room for all remaining costs and any surplus |
|---|---|---|---|---|---|
| 10m input pages; continuously rented active GPU | $10,000 | $26,192.40 | $391.90 | $0 explicitly added | −$16,584.30 |
| 10m input pages; active-processing-time rental only | $10,000 | $1,833.04 | $391.90 | $0 explicitly added | $7,775.06 |
| 80m input pages; one continuously rented active GPU | $80,000 | $26,192.40 | $391.90 | $0 explicitly added | $53,415.70 |
| Same 80m; add one passive hot spare | $80,000 | $52,384.80 | $391.90 | $0 explicitly added | $27,223.30 |
| Same 80m; add average 0.05 seconds of operator-paid review per input | $80,000 | $26,192.40 | $391.90 | $44,444.44 | $8,971.26 |
| Same 80m; both hot spare and added review | $80,000 | $52,384.80 | $391.90 | $44,444.44 | −$17,221.14 |
All rows use a scenario sale price of $0.001/input page and complete collection. $0 explicitly added means those unmeasured costs are still inside the remaining-cost question; it does not assert that no human work is required. The 0.05-second average is an assumed thirty-second intervention for every 600 inputs at $40/hour. It is not a measured error rate. The hot spare is a chosen availability stress, not a mandatory architecture or an asserted service-level guarantee.
Conditional annual budget · not a margin or a forecast
All the uncounted costs must still fit inside the bar
Scroll within the figure on a narrow screen to inspect the full chart.
Figure 1. The operator's remaining-cost budget, not its observed margin. All four plotted rows have the same modeled volume and price. The changed resources are paid explicitly. The timed batch is historical; the rental offer was inspected October 6, 2026. No real customer book or full model-development bill has been obtained.
The low-volume pair illustrates a real choice in the economics, not a choice between a favorable and an unfavorable label. An offline team that can start and stop capacity may avoid much idle rent. Cold starts, model transfer, retained disks, queued-work latency and other charges still have to be paid. A vendor promising persistent readiness cannot simply claim the active-only price without funding the readiness. Conversely, forcing an always-on dedicated fleet onto every batch buyer would exaggerate the cost.
6.4 Capital renewal and financing are not free—or counted twice
The primary calculation rents infrastructure. Its rental bill is the capacity provider's revenue and the operator's cost. Hardware and facility capital paid by that supplier must not also be deducted as though the operator bought the same assets. Extra operator-owned infrastructure, data and software investments are different costs and remain chargeable.
For an owner of hardware, the consistent alternative would replace rent with cash operations plus a justified finite capital-recovery and replacement account, including the financing convention chosen. One can model the purchase and later residual in a discounted cash flow, or an equivalent annual capital charge, but not both for the same capital. Loan principal is a financing flow, not a second hardware purchase. No matched purchase, maintenance and net resale record was obtained here, so this chapter does not invent one to produce a preferred ownership return.
This distinction also limits generalization. A cloud supplier can rent an already-installed GPU profitably on incremental terms after its original investor earned too little. A new purchaser may require a different price to recover a new investment. The observed rental offer alone cannot distinguish those cases. The older-capacity investigation's selective continued-use finding survives; it does not provide a GPU residual appraisal for this workload.[14]
Annual positive room does not establish payment timing. A small vendor may owe rent or payroll before customer invoices are collected. An advance is received once and retains its service obligation; a credit reduces receipts once. The scenario's net collection parameter is not a matched intraperiod bank account. The payer-funding and completion-to-cash investigations’ timing and availability discipline applies to the interpretation, not as evidence of a financing relationship between these parties and their earlier construction cases.[16]
7. When cheaper work helps both sides—and when volume cannot repair the price
7.1 A serious shared-gain path
Begin with 64 million inputs at $0.001, one continuously rented active GPU, and the same priced development subset. The budget for other costs and surplus is $37,415.70. Now suppose adequate paid demand grows to 80 million, while the price falls 10% to $0.0009. Receipts become $72,000 and the remaining-cost budget becomes $45,415.70: an $8,000 increase.
The customer receives a lower unit price, and the operator has more room to fund other costs. Both volumes fit the stated active capacity. This is a coherent favorable mechanism because the scenario uses spare productive time rather than assigning unlimited work to fixed resources. It does not require an arbitrarily fixed market size or assume that technical efficiency harms total profit.
The condition is precise: additional review, support, storage, development and other incremental costs caused by the expansion must fit inside the extra $8,000 before complete surplus rises. The 25% volume growth is an assumption, not a measured price elasticity or proof that customers will buy the extra work. If the additional pages are harder, the same measured throughput may not apply. A larger sales number without a quality and resource match is not the favorable case tested here.
7.2 A price reduction that reaches a paid-capacity step
Start instead from the main 80-million-page, $80,000-receipt case. A 30% price reduction to $0.0007 requires 114.286 million pages merely to preserve revenue. That exceeds the single-active-GPU scenario limit of 107.168 million. At that limit, revenue reaches only $75,017.67, leaving $48,433.37 before the remaining costs—less than the original budget, despite more work.
Adding a second continuously rented active GPU changes the expense account. To preserve the original $53,415.70 budget after paying for two active GPUs and the same priced development subset requires:
N = ($53,415.70 + $52,384.80 + $391.90) / $0.0007
N = 151.703 million input pages/year, approximately 89.63% above the original volume.
That fits the assumed two-active-GPU capacity, but the volume must actually be sold and delivered with adequate quality. Moreover, the calculation holds other costs unchanged. Additional staffing, review, storage and support would raise the volume or price needed. This is not a probability of failure. It demonstrates that preserving revenue is a weaker requirement than preserving the amount left to fund the business when expansion crosses a resource boundary.
A better validated implementation, a different workload mix, less standby, a more flexible rental arrangement or a defensible higher price may change the result. Reducing the chosen 25% capacity cushion is also mathematically possible, but not evidence that reliability permits it. Competition is an economic constraint on the receipt; neither a benchmark win nor a future demand slogan automatically finances the required capacity response.
Accepted-work economics · evidence and assumptions distinguished
Preserving revenue is not preserving room to run the business
On a narrow screen, scroll within the figure to inspect the labels.
7.3 Engineering efficiency can matter without another large training run
The development account attributes a decline in page-level retries from roughly 40% to 1% to changing the output representation. Its separately reported very low final runtime-failure rate is not a content-accuracy rate.[1]
A bounded calculation clarifies the mechanism. If each affected page incurs exactly one additional attempt of equal duration, total work per initial page changes from 1 + 0.40 = 1.40 to 1 + 0.01 = 1.01. Work falls 27.86%, or capacity per unit of work rises 38.61%. Actual retries may be shorter, longer or repeated; the reported page-level frequencies do not establish independent per-attempt failure probabilities. A geometric-retry model would impose an unsupported process.
This factor is not applied again to the measured serving batch. That would risk counting an incorporated improvement twice. It is a separate explanation of how reliability engineering can improve both quality and cost. Nor does a throughput improvement instantly lower a fixed annual rental payment. It may create saleable capacity, permit shorter rentals, reduce required nodes or absorb a harder workload. The benefit depends on the operating arrangement and demand.
The finding makes the favorable efficiency argument more concrete. Better economics need not come only from a cheaper chip or a larger model. Avoiding waste in the system surrounding the model can matter. It also leaves room for differentiated suppliers: an operator who reliably handles difficult documents can create value beyond access to common weights.
8. A materially different workload: the price of an agent is not the price of a draft
For interactive coding, an accepted unit might be a reviewed change that passes required tests and is integrated into the maintained system. A session ending successfully, a suggested line and a merged change are not interchangeable outputs. Repository context, tool execution, tests, human review and security requirements affect both usefulness and expense. A batch OCR page-rate calibration cannot supply their marginal cost or accepted-task yield.
GitHub's April 27, 2026 announcement described a move from counting premium requests to credits reflecting input, output and cached tokens, starting June 1. Its stated explanation was that long, multi-step agent sessions had made the old uniform request treatment unsustainable. That is an attributed company explanation for changing the product's economics, not an audited disclosure of its margin or proof that every individual request lost money.[10]
The operative organization documentation inspected at this cutoff supplies a firmer commercial constraint: one AI credit corresponds to $0.01 of billing value, and a standard Business license contributes 1,900 monthly credits to the organization's pool. Unused credits do not roll forward; additional usage can be paid for or blocked under budget policy. The introductory higher allocation ended September 1. Code completions remain outside credit billing. These are entitlements and billing rules, not counts of accepted pull requests or measured provider compute costs.[11]
The simple calculation 1,900 × $0.01 = $19 identifies a billing allowance, not nineteen dollars of observed inference expense. The product can include some low-cost, high-frequency functionality while separately metering more resource-intensive behavior. A constant base subscription price therefore does not imply a constant amount of all useful work, and unlimited completions do not imply unlimited autonomous agent work.
The consequences differ from the document operator's batch case. Interactive work cannot always be delayed to fill a discounted queue. Tool execution and review remain after a model call; the announcement also separates Actions-minute charges for code review from model credits. Historical annual-plan exceptions mean a transition date should not be treated as proof that every individual contract immediately changed.[10]
The favorable interpretation is a provider adapting prices and budgets so that heavy useful work can finance its costs instead of being indiscriminately rationed. The adverse interpretation is that a buyer's desired workload becomes more expensive, a budget limit interrupts it, or metering makes a cheaper adequate alternative attractive. Neither follows solely from the announcement. The observable advance is the change in which usage consumes the paid allowance, not a finding of universal subsidy, a coding productivity estimate or a stock valuation.
This contrast limits a sweeping claim that falling model prices must preserve all AI product margins. A service can use cheaper components while users ask it to perform more steps, retain more context or execute longer processes. Pricing and inclusion rules can change who pays for that extra work. The user, the procurement authority and the gain recipient again need to be identified separately.
9. Who retains the gain when adequate alternatives exist?
9.1 The buyer and workflow operator
A buyer can retain a lower external bill, better service, less review or more accepted work. The application operator can retain the difference between a stable workflow price and lower component costs, but only while competition, customer bargaining and switching terms permit. A pure pass-through extraction seller faces a different competitive position from an application offering tested integrations, provenance and accountable handling of exceptions.
This does not prove an incumbent's pricing power. A buyer may be able to migrate a narrow component while keeping its own workflow. Another may face substantial data-security, audit and integration work. Those are real costs, but not a license to assume permanent lock-in. The cost allowance in §5 is useful precisely because it makes a migration proposal falsifiable: measure the extra work and compare it with the amount saved at the relevant volume.
9.2 The model developer and host
A released open model can lower a user's licensing or serving bill while relying on development paid elsewhere. A paid API can charge for managed availability, changing models and additional functionality. These are different funding arrangements, neither inherently fictitious. The investigation does not establish Ai2's complete cost recovery, Mistral's net cash margin or the cloud supplier's return on owned capital.
Distribution also matters. The inspected DeepInfra page for the 1025 olmOCR model displays a low-usage deprecation notice while also saying requests are still working. Its numeric date string, 05/7/2026, is ambiguous and the served page is cached. No endpoint was tested; no completed shutdown is asserted. The supported observation is a host's displayed notice about that listing, not a measurement of worldwide OCR demand.[9]
The same page offers token pricing for raw hosted inference. That is not substituted into the page-cost calculation: its token usage, model version, surrounding toolkit, request terms and current availability are not matched to the timed batch. The notice illustrates a practical point without overclaiming it: open weights can remain available while one convenient commercial access route changes. Keeping the service useful may require migration, deployment work or another supplier, all of which consume part of the gain.
9.3 The capital owner
Lower serving costs can increase the profitable uses of a capacity pool, but may also lower its scarcity rent. Higher physical throughput helps only when adequate paid work occupies it or resources can be reduced. An existing asset may keep generating positive incremental cash without a new copy being worth its full purchase price. A model developer may create broad customer value without capturing a large percentage of it.
These are not contradictory outcomes. A technical improvement changes the possible combinations of output and resources. Contracts, competition, distribution, financing and ownership prices determine how much of that improvement each party retains. The report quantifies one such combination and an alternative billing response. It does not equate gross customer benefit with supplier revenue, or add rental receipts to the operator's sales as independent final demand.
10. What changes in the whole answer—and what remains conditional
The productive-efficiency case is strengthened. The evidence now includes a documented conversion pipeline, a historical downstream comparison, reported serving time and an operative metered-service response. It is no longer necessary to explain efficiency only through an abstract fall in token price. There are identifiable ways to deliver useful work with modest machine expenditure and to alter pricing so heavy work is not treated as free.
The claim of effortless full-cost capture is weakened. A large percentage component saving can be small in dollars per page. In the principal model, less than a tenth of a second of additional average checking can erase the active-time saving over a cheap API. A vendor's annual residual depends strongly on shared versus dedicated capacity and on whether support and intervention stay automated. The calculated residual is valuable because it states the maximum remaining-cost envelope; it does not turn unmeasured full development into zero.
Demand expansion remains a conditional, potentially favorable mechanism. The 64m-to-80m path permits a lower price and a larger remaining-cost budget. The 30% price-cut path shows why preserving revenue can be insufficient once another machine is required. Neither is an estimated demand elasticity. The visible price list and technical capacity do not establish paid uptake.
The accumulated reports are revised in meaning, not silently recalculated:
| Prior finding or assumption | Consequence of this investigation | Boundary retained |
|---|---|---|
| The ownership-price investigation's frontier cases require complete owner-cash margins, including continuing development and reinvestment. | There is concrete evidence that selected useful work can have cheap serving and that resource metering can improve cost recovery. There is also a quantified reason that serving economics alone cannot supply the full margin. | The earlier 20–30% complete cash-margin assumptions remain assumptions; none is replaced by the document rental spread or H/revenue. |
| The payment-and-capture investigation finds that buyers and providers can both retain value. | A finite shared-gain path now makes the condition explicit: incremental unpriced costs must fit within the extra budget after the lower price and paid capacity. | Company-level retention and expense accounts remain different from a model-service cohort or this hypothetical operator. |
| The older-capacity investigation shows how old capacity can remain useful without earning the original return. | The rental/ownership distinction and specialization show how a service can remain viable under another deployment arrangement. | No old hardware tariff ratio or Austin expense budget becomes a current GPU appraisal; original recovery is not inferred. |
| The buyer-budget investigation distinguishes accepted benefit from new expenditure. | Lower extraction costs may free existing budgets or support more useful processing; only part need become supplier receipts. | Council assessments and marketing assets are not the measured OCR workload, nor a calibration of its acceptance or review. |
| The completion-to-cash investigation distinguishes available resources from eventual support. | Flexible rentals and metered service can reduce a future burden, but annual operating room still needs actual collection and payment timing. | No causal or contractual connection to the completion-to-cash investigation's facilities is invented, and its conditional reserve result is unchanged. |
Sources for the inherited findings are the earlier dated investigations, not a new audit of their financial originals.[12][13][14][15][16]
Weighted conclusion
The evidence is most supportive of selective economically useful expansion with contested capture of the gain. It is less supportive of either a uniformly uneconomic AI service layer or an inference that technical efficiency validates the whole build-out's ownership prices. A service user has reason to distinguish the survival of a useful workflow from the survival of a particular access route and from the return to an investor who paid a high entry price.
This is not a universal bubble verdict. The selected workload favors specialization and batching; interactive coding demonstrates a different resource and billing response. Neither measures the prevalence of sustainable margins across frontier labs, hosting firms or applications. There is no sector-wide loss estimate, crash calendar or personal trading recommendation.
The strongest evidence that would make the assessment more favorable is a matched commercial record of adequate accepted work, net paid receipts, limited review/fallback, repeat demand and full continuing costs within the calculated envelope. Evidence of paid volume filling existing capacity would strengthen the shared-gain mechanism. Evidence of repeat customers buying higher-priced provenance or reliability because measured downstream work falls would support durable differentiation.
The strongest evidence in the other direction would be persistent extra checking, repairs or reliability resources that consume the apparent saving; paid demand too small to spread standing costs; price concessions that require capacity the receipts cannot fund; or ongoing development expenses beyond the residual despite adequate serving. Those observations could establish a weak service business without implying that the output has no customer value or that every creditor suffers a loss.
11. Methods, access and reproducibility
Earlier dated investigations retain their original evidence boundaries. They provide the prior findings discussed here, not an independent re-audit of their original financial sources.
Public research used primary developer papers, original provider model/pricing documentation and GitHub's own announcement and operative billing documentation. No paid inference, benchmark run, private telemetry, customer contact or procurement transaction was performed. The timed run is author-reported; the current-rate cost is a calculated hybrid. Contract discounts, realized receipts and private full-cost allocations remain unavailable. No national demand estimate or probability calibration is attempted.
The ACL PDF's printed cost tables and runtime appendix were successfully rendered and visually inspected. Earlier arXiv/PDF image attempts failed; that historical source was read through its HTML text and tables. The current offers are what the accessed pages displayed, not authenticated quotes or guaranteed capacity. The DeepInfra notice has the explicit date/cache limitation described above. An inaccessible or stale source does not establish that a service stopped or a payment failed.
The CSV preserves observation, tariff, estimate, conflict and assumption status separately. The Python companion reconstructs costs from time and rate; prices the explicit development subset; enforces finite active capacity; separates standby from throughput; includes failed inputs in the accepted-corpus illustration; and reports the remaining-cost budgets. Conflict-tagged table values are blocked from ordinary calculation access. No observed benchmark score is converted into acceptance probability.
The primary equations use a one-year steady-service comparison, not a startup cash ledger or a company valuation. A startup must separately finance deployment and the interval before net collections; future years require their own prices, demand, updates and resource availability. The figure uses the same CSV/code and identifies its missing costs. Executable checks validate the calculations; they do not validate future assumptions or replace an experimental replication.
Sources and inspection boundaries
Source 1
Jake Poznanski, Kyle Lo and Luca Soldaini, The OLMOCR Project: Building Fully Open OCR using VLMs, ACL 2026 System Demonstrations, July 2–7, 2026, pp626–635. Original paper. Key locators: §2.3 scoring; §3 development; Table1/§5 printed cost denominator; §6 retries; AppendixB training; AppendixE timing and historical prices. Tables on printed pp630 and635 were visually checked. Author-reported experiments, not independent production telemetry. The disagreement between the printed denominator and timed-cost calculation is preserved and explained in §4, not silently repaired.
Source 2
Poznanski and colleagues, olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models, version2, June18,2025. Versioned HTML. §2.3 training-run versus full experiment time; §4.2/Table5 downstream same-corpus comparison; AppendixB timed hardware and historical GPT-4o billing comparison. The old experiment is a separate model/workload vintage. HTML inspection succeeded; a complete visual PDF audit did not. Its API prices are not current offers.
Source 3
Ai2, olmOCR 2: Unit test rewards for document OCR, October22,2025. Developer release. Deployment-cost statement and release/toolkit information. Developer-authored and a different release record from the later ACL paper; used to corroborate cost scale, not add another independent sample or update the exact timed run.
Return to citation: ↩ 1
Source 4
Ai2, olmOCR-2-7B-1025 model card. Original developer model page. Released weights, license identification and surrounding toolkit requirements. The 1025 card is not a fresh timing test of the later paper's listed 1125 result. No model was downloaded or run.
Source 5
RunPod, GPU Cloud Pricing, Pods GPU table and separate storage/cluster sections. Pricing page, inspected October6,2026; served page carries a July27,2026 update label. H100SXM80GB $2.99/hour; B200 $5.89/hour; L40S $0.99/hour. These are displayed Pod offers, not the different serverless or reserved-cluster terms, actual invoice costs or a guarantee of original-test reproducibility. No checkout or capacity reservation was attempted.
Source 6
Mistral, OCR 3, model card v25.12, December18,2025. Original documentation. Ordinary-page pricing, separately annotated-page pricing and continued availability for existing integrations. Inspected October6,2026. No claim of realized average pricing or equivalence to every newer service.
Source 7
Mistral, Mistral OCR 4: SOTA OCR for Document Intelligence, June23,2026. Original release and service definition. Pricing/availability, extraction-versus-DocumentAI distinction, structural output, use boundaries and qualification of internal competitor reproductions. Provider claims were not independently replicated. The calculation uses the ordinary extraction offer, not an assumed all-inclusive annotation or decision service.
Source 8
Mistral, Batch Processing. Operative documentation, inspected October6,2026, undated page. The 50% price discount concerns asynchronous processing. No real customer discount, queue time, successful-job rate or low-latency service guarantee is inferred.
Source 9
DeepInfra, allenai/olmOCR-2-7B-1025 model listing. Displayed listing, inspected October6,2026. Notice attributes deprecation to low usage and displays 05/7/2026 while saying requests still work; the served page is cached. This does not verify shutdown timing, present endpoint status, global demand or profitability. Its token offers are not used as matched page costs.
Return to citation: ↩ 1
Source 10
GitHub, Mario Rodriguez, GitHub Copilot is moving to usage-based billing, April27,2026. Original announcement. June1 transition, stated cost rationale, included features, code-review Actions-minute distinction and legacy annual-plan exceptions. The explanation of unsustainability is attributed to GitHub, not an audited margin finding. Later operative terms are sourced separately.
Source 11
GitHub, Usage-based billing for organizations and enterprises. Operative documentation, inspected October6,2026. Billing-value conversion, standard monthly pooled allowances, September1 end of the promotional period, no rollover and budget/additional-usage treatment. These rules identify entitlement and payer exposure, not observed token costs or completed useful work. No user account, policy or budget was accessed or changed.
Source 12
What today’s ownership prices require, evidence boundary 26 September 2026, especially the frontier-ownership section. Its complete owner-cash margins and funding choices remain conditional; neither is recalibrated with an extraction spread.
Return to citation: ↩ 1
Source 13
Paid persistence is real; the division of the gain is not uniform, evidence boundary 29 September 2026. Full expense versus cash and company-versus-product distinctions remain. The signed working-capital result is unchanged; no Freshworks margin enters this model.
Return to citation: ↩ 1
Source 14
Older capacity can earn again—but it does not become free, evidence boundary 29 September 2026, historical Austin anchor 31 March 2025. Source tariffs, older-design contracting, expense-derived room and original investment recovery remain different claims. No prior hardware-price ratio supplies the present workload’s cost result.
Source 15
Buyer demand and budget substitution, evidence boundary 5 October 2026. Council purchased capacity, reported outcomes and marketing-cost attribution are not accepted-page or serving-cost observations for the workload examined here.
Source 16
The earlier demand investigation, payer-funding investigation, holder-exposure investigation, joint-cash investigation and completion-to-cash investigation retain their individual evidence dates. NVIDIA–Energy Global: obligation established, receipt unverified, no nonpayment inference. The later CoreWeave closing remains a dated financing, not current unspent cash. The completion investigation resolved April’s 8.25% issuer illustration versus May’s 7.75% executed coupon; this is not a new correction or a rerun of the older model. These cash and source-status distinctions inform interpretation without establishing a relationship to the document workload.