Edition 22 · Weekly · APAC · Infrastructure Finance

Where Inference Has to Run

The most valuable megawatt may not be the closest one. It may be the least substitutable one.

By Sel Fang, Lim · Data centre and infrastructure finance, APAC
Subscribe on LinkedIn Read the editions →
10% uplift
OpenAI & Mistral geographic processing surcharge
21 → 93 GW
McKinsey inference demand, 2025–2030
4 workload patterns
Real-time / sovereignty / data gravity / latency-tolerant

Edition 21 asked what happens when AI starts to change what "location" even means. This is that edition. Inference doesn't migrate to one new geography — different workloads have very different answers to where they're actually allowed, or economically sensible, to run.

AI infrastructure has spent the past several years chasing power. Large training campuses made the logic easy to understand: secure enormous power blocks, concentrate accelerators, build dense networking, optimise the factory around throughput. Inference complicates that geography — not because inference simply "moves to the edge." It doesn't. The more important change is that different inference workloads have very different answers to a more basic question: where are they actually allowed — or economically sensible — to run?

1 The deal

That distinction matters to an investor because the cheapest place to build compute is not necessarily a place where every workload can actually run. Each workload has its own placement envelope — the set of locations that can satisfy its technical, regulatory and economic requirements. The wider that envelope, the more locations can substitute for one another. The narrower it becomes, the scarcer an acceptable location can become. A location premium only has a chance to exist where that scarcity starts to bite — but scarcity alone does not guarantee an owner captures it, a distinction this edition returns to throughout.

Geographic processing now has a printed price — but it's a usage fee, not a rent. OpenAI and Mistral each explicitly charge a 10% uplift for eligible regional-processing / regional-inference endpoints (both vendor documentation, retrieved August 2026). That does not establish a 10% data-centre location premium. The surcharge accrues at the model-service layer. The infrastructure question is how much, if any, survives into facility utilisation, contract durability or rent — and that is the question this edition works through.

It matters because inference is no longer the smaller half of AI compute. McKinsey models inference-related data-centre demand rising from roughly 21 GW in 2025 to 93 GW by 2030, against training demand rising from about 23 GW to 62 GW over the same period — though McKinsey's own modelling assumes AI data centres can serve both workloads, so this is not a clean physical split. Inference is projected to overtake training during the forecast period and non-AI workloads by 2029.

2 The engineering read

First, separate two different latencies

One reason the inference debate gets confused is that "latency" is treated as a single thing. It isn't.

Intra-cluster latency: inside an AI factory, accelerators exchange enormous volumes of data at extremely low latency — high-bandwidth, low-latency GPU-to-GPU communication is central to how modern training and inference workloads (including mixture-of-experts models) actually run. Training should not be described as "latency insensitive." It may be relatively insensitive to its distance from the eventual end user while being extraordinarily sensitive to communication inside the compute cluster.

Location latency: a different question — the time required for inference infrastructure to communicate with the user, enterprise application, private data source, or another system the model repeatedly calls. For interactive AI, that can affect user experience. For machine-to-machine or time-critical applications, the constraint can instead be operational: the application itself may have a tight latency budget even when no human is waiting for the response. AWS's Amazon EC2 G7e instances — accelerated by NVIDIA RTX PRO 6000 Blackwell GPUs — became available in Asia Pacific (Tokyo) in February 2026 and Asia Pacific (Seoul) in March 2026. In July 2026, AWS extended G7e support to SageMaker AI inference in both markets, explicitly framing the expansion as allowing inference endpoints to sit closer to Asian end users and reduce generative-AI latency. That is evidence proximity can matter. It is not evidence all inference needs to be local.

Inference is not one workload. Four different workload patterns can produce very different placement envelopes:

Real-time & time-critical (proximity-sensitive) · Sovereign/regulated (in-border, eligibility) · Enterprise (data gravity) · Latency-tolerant (pooling, utilisation, power economics)

Real-time and time-critical inference

This covers two related but distinct cases. Human-facing: voice, video, interactive generative AI, customer-facing applications, where response time affects user experience directly. Machine-facing: applications or systems with tight operational response-time requirements, where location latency can matter even without a human directly waiting for the response. Tokyo or Seoul inference capacity can have a real advantage here even where cheaper accelerator capacity exists farther away.

Sovereign or regulated inference

The relevant boundary isn't milliseconds, it's a national border. Unlike latency or data gravity, this is fundamentally an eligibility constraint: even if another region is cheaper, faster or has more available compute, it remains outside the placement envelope if the workload cannot legally or contractually be processed there. AWS introduced India Geographic cross-region inference in August 2026, initially for OpenAI models on Bedrock, letting customers with in-country processing requirements route requests between Mumbai and Hyderabad while inference stays inside India. Sovereignty does not necessarily eliminate pooling — it changes the boundary within which pooling can occur. The workload is geographically constrained but not necessarily metro-constrained — sovereignty creates a placement envelope without creating an edge requirement. Japan's Digital Agency is running a parallel version: a fiscal-2026 pilot testing three domestically developed foundation models on Sakura Cloud — the Government Cloud's only domestically developed cloud service — with blind evaluation trials running September–November 2026 alongside existing access to Amazon Nova and Anthropic Claude models. The constraint in both cases is eligibility, not latency.

Enterprise inference — data gravity

A different kind of gravity, created by repeated interaction with private databases, applications, security controls and other enterprise systems. Here, the placement constraint is not necessarily end-user latency. Moving inference farther away can increase network dependency, data movement, response time and architectural complexity when the model repeatedly needs to retrieve, process or act on enterprise data. In those cases, compute may be pulled toward the data and systems it depends on — rather than toward the end user.

What Equinix's Inference Exchange is — and isn't

Equinix announced its Inference Exchange on 2 September 2026, combining NVIDIA Enterprise Reference Architectures with Together AI's inference platform to connect distributed inference infrastructure to enterprise data, clouds and networks — but the service is announced for availability from Q1 2027, with no pricing, capacity commitments, or anchor customers yet disclosed.

It's useful evidence of where the industry is investing. It is not yet proof that enterprises will pay data-centre owners a measurable premium for it.

Latency-tolerant inference

Batch workloads, asynchronous processing, background AI tasks. These have much weaker reasons to sit close to the user. The priority instead becomes GPU utilisation, power cost, available capacity, and cost per token. Proximity has little economic value where latency is non-binding, which can pull compute away from expensive metros rather than toward them.

The counterforce: not all inference demand needs local inference capacity

Remote inference itself is not new. A Singapore application could already call compute hosted in another region, and companies could build applications across multiple regions themselves. What is becoming more explicit is who manages the placement. Instead of a customer having to decide where inference capacity is deployed and manually select which region should serve an eligible workload, that placement can increasingly be managed automatically at the managed-model service layer. The service can dynamically draw on model capacity across a broader pool of supported regions.

AWS's Global cross-Region inference makes this visible. Customers in Thailand, Malaysia, Singapore, Indonesia and Taiwan can invoke supported Claude models while AWS's managed inference service dynamically routes eligible requests across 20+ supported commercial regions worldwide. The customer does not have to manually select the serving region for each eligible request. That does not create remote inference. It makes cross-region placement increasingly productised and automatically managed at the model-service layer.

The infrastructure implication is important. If a workload does not have to stay in Singapore, for example, Singapore demand does not automatically require Singapore inference capacity. In other words: demand originating in a market ≠ compute that must be hosted in that market. For the data-centre investor, this changes the scarcity equation. Cloud pooling expands the set of locations that can serve a workload. Latency, sovereignty, data gravity and other constraints shrink it.

This is where a simple "inference moves to the edge" thesis breaks. Inference carries two opposing forces at once: cloud pooling can make capacity across regions increasingly interchangeable for workloads that do not need to stay local, while proximity, sovereignty and data gravity can pin other workloads to a much narrower set of locations. The investment question is therefore not whether inference is becoming local or global. It is which workloads can be pooled — and which ones cannot.

The seam worth marking. A geographic constraint can narrow the eligible pool without eliminating pooling inside that boundary. The more important investment question is whether that constraint creates value the facility owner actually captures, rather than value retained by the hyperscaler, network provider or GPU cloud.

3 The capital allocation read

1. Geographic processing has a printed price — but it's a usage fee, not a rent.

OpenAI and Mistral both publish a 10% uplift for eligible geographically constrained processing/inference offerings — the market pricing "here, not there" explicitly. (Azure OpenAI and Google Vertex AI also offer region-scoped deployments; a consistent, directly comparable price differential was not confirmed against their own documentation, so no claim is made about them either way.) That surcharge is charged per API call to the model vendor. It is not, today, an offtake payment to the facility hosting the endpoint. A usage fee only becomes an owner's return if it survives the path into facility economics. Otherwise, the premium belongs to another layer of the stack — not the data-centre investor. The underwriting question is never "this facility is close to users." It is: what scarce function does the facility control that the customer cannot cheaply reproduce elsewhere?

2. Workload eligibility — not market — is the unit of underwriting.

The width of that placement envelope determines how many viable substitutes a workload has. A workload with twenty acceptable locations gives the customer bargaining power. A workload with one is where location becomes strategic. Conceptually, location scarcity rises as the number of viable substitutes falls. Managed cross-region inference can expand that placement envelope for eligible workloads because the serving location can increasingly be selected dynamically at the managed-model service layer rather than manually tied to the market where demand originates. But that expansion stops wherever a workload constraint binds — latency may narrow the pool, sovereignty may confine it to one country, enterprise data gravity may pull it toward a particular ecosystem, service availability or application architecture can narrow it again. So the relevant question is not how many cloud regions exist. It is how many of them are actually viable substitutes for this particular workload.

Malaysia is increasingly part of the same cloud architecture that makes Singapore valuable — AWS operates separate Singapore and Malaysia regions, and Microsoft's Malaysia West (Kuala Lumpur) region has been generally available since May 2025, with a second Malaysian region in Johor Bahru ("Southeast Asia 3") announced by Microsoft but not yet given a confirmed launch date. AWS already lets eligible inference invoked from both Singapore and Malaysia draw on globally distributed model capacity through its managed cross-region inference architecture. That makes "can inference move from Singapore to Johor?" the wrong question. The better one: which workloads still require characteristics that Singapore provides once Malaysia itself becomes a genuine cloud and AI region? If a workload's latency, residency and connectivity requirements are satisfiable in Malaysia, the two markets become more substitutable for that workload. If not, Singapore retains scarcity. That is a different investment proposition from simply "Singapore is closer to demand" — and this edition is not claiming a Singapore-versus-Johor IRR ranking either way.

3. The sovereignty constraint may be the more durable one.

A latency-driven placement requirement can erode as networks improve and as on-device silicon absorbs the short-call tier. A legal in-country requirement doesn't erode simply because fibre improves — it persists until the regulatory or compliance boundary changes. That makes sovereignty a potentially more persistent constraint on where a workload can run. Persistence of the constraint is not the same as proof of an owner-side premium. Owner capture is a function of location scarcity, infrastructure control, and contractual retention together; if another layer of the stack holds the control point, the facility owner may see very little of it regardless of how scarce the location is.

Placement-envelope underwriting pass

  1. What is this workload's placement envelope — how many locations satisfy its latency, residency, connectivity, reliability and economic requirements?
  2. Is scarcity driven by latency, sovereignty, or data gravity — or some combination?
  3. Does pooling economics outweigh locality for this workload — can a managed service satisfy it from elsewhere in an eligible pool?
  4. What scarce function does the facility actually control that the customer cannot cheaply reproduce elsewhere?
  5. Does that control survive into rent, utilisation or contractual durability — or does another layer of the stack capture it first?

Investment Lens

Constraint
Location latency (real-time and time-critical inference)
Impact
Response time can constrain either user experience or the operational latency budget of a machine-facing application
Capital response
Metro/near-metro inference endpoints where proximity materially improves the workload's response-time requirement (e.g. AWS G7e in Tokyo, Seoul)
Candidate beneficiary
Metro colo supporting latency-constrained workloads — placement advantage depends on how tightly proximity binds
Constraint
Sovereign or regulated inference
Impact
A national border, not a millisecond budget, defines eligibility
Capital response
In-country compliant endpoints (India Geo CRIS, Japan's Sakura Cloud pilot)
Candidate beneficiary
Infrastructure or service layer controlling compliant in-country compute — capture remains unproven
Constraint
Enterprise data gravity
Impact
Repeated interaction with private data, applications and enterprise systems can make some locations or ecosystems more practical than others
Capital response
Interconnected inference infrastructure positioned close to enterprise data, clouds and networks
Candidate beneficiary
Interconnection / colo / service layer controlling access to the relevant enterprise ecosystem — owner-side capture remains unproven
Constraint
Latency-tolerant / pooled inference
Impact
Eligible requests can draw on a broader regional model-capacity pool; the originating market does not necessarily determine where inference runs
Capital response
Managed cross-region routing across a wider pool of supported capacity (AWS Global CRIS)
Candidate beneficiary
Power-rich, scalable campuses able to participate efficiently in large shared inference pools
Constraint
Value capture (usage fee → facility rent)
Impact
The published 10% geographic-processing uplift (OpenAI, Mistral) accrues at the model-service layer, not directly to facilities
Capital response
Contract for endpoint control or residency compliance, not just space and power
Candidate beneficiary
The entity holding the compliant, non-substitutable endpoint — capture not yet demonstrated

4 What this means for the broader market

The AI infrastructure build-out began with a power question: where can we build the next large block of compute? Inference adds another one: where can this particular workload actually run? Those are not always the same place. Some inference will remain geographically flexible, with managed-model services able to route eligible requests automatically across broader regional capacity pools. Some will remain inside national borders. Some will cluster around enterprises and cloud ecosystems. Some will genuinely move closer to users or time-critical systems. Location value therefore comes from disappearing substitutes, not from proximity alone — and value only reaches the facility owner if the owner, not another layer of the stack, controls the bottleneck.

On-device inference is another live counterforce: as capable local silicon absorbs some latency-sensitive workloads, the pool of workloads that actually requires metro data-centre inference could shrink.

The investment opportunity, then, is not simply to own "edge" capacity. It is to identify where a growing workload faces a shrinking set of acceptable locations — and then determine whether the infrastructure owner actually controls that bottleneck.

The most valuable megawatt may not be the closest one. It may be the least substitutable one. And that raises the next question — if a durable constraint doesn't automatically hand its premium to whoever sits inside the border, is the winning asset a data centre at all, or the control point wrapped around it?

That is the question I'm taking forward next.

Sources & research notes

This edition draws on vendor pricing disclosures, primary AWS/Microsoft/Equinix announcements, government-agency releases, and tier-1 market research current to early September 2026. Figures are tiered; directional and announced-not-operational items are flagged.

Vendor pricing (geographically constrained processing) — OpenAI: published 10% uplift for regional-processing (data residency) endpoints on eligible models, retrieved August 2026. Mistral: regional inference (EU/US endpoints) billed at 1.1× standard list pricing, per Mistral's own documentation, retrieved August 2026. Azure OpenAI and Google Vertex AI offer region-scoped deployment options; a directly comparable, confirmed price differential was not established against their own primary pricing pages and is not asserted here.
AWS — Amazon EC2 G7e instances: general availability in Asia Pacific (Tokyo, February 2026) and Asia Pacific (Seoul, March 2026); G7e support extended to SageMaker AI inference in both regions (July 2026), with AWS's own announcement framing the expansion around lower generative-AI latency for Asian end users; India Geographic cross-region inference for OpenAI models on Bedrock, Mumbai/Hyderabad (August 2026); Global cross-Region inference for Claude models in Thailand, Malaysia, Singapore, Indonesia, Taiwan (February 2026). For Global cross-Region inference, the relevant architectural point is not that remote inference is new. It is that eligible serving-region selection can be dynamically managed at the managed-model service layer across a broad supported capacity pool rather than requiring the customer to manually select the serving region for each request.
Microsoft — Malaysia West (Kuala Lumpur) region live since May 2025 (Microsoft Asia newsroom); second Malaysian region "Southeast Asia 3" (Johor Bahru) announced, planned, no confirmed launch date as of this writing.
Equinix — Inference Exchange (with NVIDIA, Together AI) announced 2 September 2026; availability from Q1 2027; no pricing or anchor customers disclosed at announcement. [ANNOUNCED — not yet operational]
McKinsey — inference-related data-centre demand ~20.9 GW (2025) to ~93.3 GW (2030), 35% CAGR; training ~23.1 GW to ~62.2 GW, 22% CAGR; McKinsey's own modelling assumes AI data centres serve both workloads (not a clean physical split); inference moves ahead of training during the forecast period and overtakes non-AI demand by 2029.
Japan Digital Agency — Government AI GENAI pilot testing three domestic foundation models (NTT tsuzumi 2, Fujitsu Takane 32B, Preferred Networks PLaMo 2.0 Prime) on Sakura Cloud; blind evaluation trials September–November 2026; feeds FY2027 procurement decisions. [PILOT — not full deployment]
APAC data-residency analysis (2026) — inference on locally-originated data via a foreign endpoint may count as a cross-border transfer under APPI (Japan), PIPL (China), PDPA (Singapore).
Verification notes
[VERIFIED] — Only vendor pricing directly confirmed against primary documentation is cited (OpenAI, Mistral).
[MONITOR] — Azure OpenAI and Google Vertex AI region-scoped pricing: a consistent, directly comparable differential was not established against their own pricing pages. Re-verify before any future edition asserts a figure.
[ANNOUNCED] — Equinix Inference Exchange and Microsoft's Southeast Asia 3 (Johor) region are both pre-operational; treated as directional evidence of industry intent, not delivered capacity.
[PILOT] — Japan's Sakura Cloud programme is a fiscal-2026 evaluation pilot feeding FY2027 procurement decisions, not a completed deployment.
[HOUSE VIEW] — "Sovereignty constraint may be more durable than latency constraint" is the author's inference from the legal-vs-technical erosion mechanism, not an observed owner-return comparison.

Glossary

TermFull namePlain English
APACAsia-PacificThe region this newsletter tracks.
APPIAct on the Protection of Personal InformationJapan's core data-protection law.
CRISCross-Region InferenceAWS's mechanism for routing inference requests across a defined pool of regions.
GPUGraphics processing unitThe accelerator hardware used for AI training and inference.
GWGigawatt1,000 megawatts — used here for aggregate demand modelling.
LLMLarge language modelThe class of AI model this newsletter's infrastructure discussion is built around.
MWMegawattThe standard unit of data-centre power capacity.
PDPAPersonal Data Protection ActSingapore's core data-protection law.
Placement envelopeThe set of locations technically, legally and economically able to run a given workload.
PIPLPersonal Information Protection LawChina's core data-protection law.
Regional inference / regional processingA vendor product tier that constrains where a model executes to a chosen geography, typically at a price premium.
TTFTTime to first tokenA common latency metric for generative AI responsiveness.
The Uptime Brief · Where uptime meets capital allocation · By Sel Fang, Lim · Published weekly

Editorial note: This publication is provided for informational and educational purposes only. It reflects the author's analysis of publicly available information as of the publication date and should not be construed as investment, legal, accounting, or financial advice. Opinions are the author's own and may change as further disclosures become available. No investment advice intended or implied.