Future TechnologyFuture Technology
HARDWARE

Arm's biggest GPU redesign in seven generations shipped first in China

· 5 min read · By Future Technology

Key takeaways

  • The Xiaomi 18 Fold is the first shipping device with Arm's Mali G2-Ultra NX, inside Xiaomi's Xring O3 chip
  • A new matrix accelerator handles up to 1,024 INT8 operations per clock and powers Arm's Neural Super Sampling
  • It supports INT8 and INT16 only, with no BF16 or FP8, so it is built for graphics rather than general on-device AI
  • Arm's 14 percent gaming uplift is measured at roughly 11 percent higher clocks, so most of the architectural gain lands in ray tracing

The first phone shipping Arm's Mali G2-Ultra NX graphics processor is the Xiaomi 18 Fold, and it is on sale in mainland China. The GPU sits inside Xiaomi's own Xring O3 chip, which launched on 24 August 2026.

Arm describes the part as the largest re-architecting of its GPU IP in seven generations. Whatever the marketing is worth, the substance is that a British designed graphics architecture with dedicated machine learning hardware reached Chinese consumers before it reached anyone else.

What the neural hardware actually does

Arm added a matrix accelerator to the shader core, the same class of unit Nvidia and AMD have used on desktop for years. Each one handles up to 1,024 INT8 multiply-accumulate operations per clock, or 512 at INT16, and it can run at up to twice the clock of the shader core's normal execution engines.

That hardware is what enables Neural Super Sampling, Arm's version of machine learning upscaling. The GPU renders each frame at a lower internal resolution and reconstructs it, rather than paying the full cost of every pixel at output resolution. Arm pairs it with neural frame rate upscaling and neural denoising, and claims that running all three together can produce up to a fourfold increase in frame rate.

That is a ceiling, not an average. Upscaling gains depend heavily on how far below native the game renders and how much motion is on screen, and the same claim structure has been used for every desktop upscaler since DLSS arrived.

The precision support is narrower than it looks

The matrix accelerator supports INT8 and INT16 only. There is no FP8, and no BF16, which is the format most current machine learning work is trained and run in. The standard shader ALUs do not support BF16 either.

That narrowing looks deliberate. The unit is built for the graphics inference workloads Arm has already chosen, not for general on-device model work, and developers hoping to reuse the silicon for anything else will find the door mostly closed.

It also costs area. A shader core carrying the matrix unit measures roughly 1.88 square millimetres against 1.55 without it, about 21 percent more silicon. Arm requires a minimum of six NX-equipped cores for the Ultra branding but does not require every core to have one, and Xiaomi took that option: only half the shader cores in the Xring O3 carry a matrix accelerator.

Read the performance numbers with the clock speed attached

Arm claims up to 14 percent improvement in games and up to 24 percent in ray tracing benchmarks against the previous generation. The comparison runs the G2-Ultra NX at roughly 11 percent higher clocks than the G1-Ultra, which accounts for most of the gaming figure and suggests the architectural work pays off mainly in ray tracing rather than conventional rendering.

The ray tracing changes are specific and real. Arm made the triangle structure the RT unit operates on more compact, which it says cuts DRAM traffic by 13 percent by removing redundant data and letting more of it stay in cache. Opacity micromaps, another feature imported from desktop GPUs, arrive at the same time.

Xiaomi's own claims for the Xring O3 are larger: 85 percent higher GPU performance and 64 percent lower power draw than the Xring O1, with a 182 percent jump in ray tracing. Those compare an entire new chip on a newer process against an older chip, not architecture against architecture, so they are measuring a different thing.

Why shipping in China first matters

Mobile studios tune their games for the hardware that exists in players' hands. Wider commercial availability of G2-Ultra NX is expected in 2027 Android phones and larger-screen devices such as Chromebooks, which leaves roughly a year in which the only meaningful install base for Arm's neural rendering path is Chinese.

Studios building for that market get a head start on integrating Neural Super Sampling and on learning where it breaks. Everyone else will integrate it later, against hardware they cannot yet buy, and will inherit optimisation work done for a different audience on a different set of titles.

For a technology whose entire value depends on developers choosing to use it, that ordering is the part worth watching. Arm has built the hardware. Where the software gets written first decides how quickly it matters anywhere else.

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.

Enjoyed this? Get the briefing.

One email, every weekday: the top story, a useful tool, and what matters in tech - in under 3 minutes.

More from Future Technology

Security

"A Shipping Partner Breach Just Exposed Thousands of Trezor Buyers"

Security

"A Chinese Hacking Crew Turned a VMware Bug Into a Ransomware Pipeline"

Security

Home Network Security Settings for 2026: Nine Router Changes Worth Twenty Minutes

Security

Abliteration.ai Wants to Sell You Unrestricted AI Models, and the Debate Is Complicated