Xiaomi's September 7 announcement of the 18 Fold pairs a wide foldable design with the company's XRING O3 processor and several local-AI claims. Independent launch reports agree on the headline facts: the phone was announced for China with a 7.58-inch inner screen, a 5.38-inch outer screen and a 6,000mAh battery. Those facts make the device interesting. They do not, by themselves, prove that local AI is private, useful, fast under load or well supported by everyday software.

This distinction matters whenever a phone maker presents a new neural processor, memory standard or benchmark score. A specification sheet describes hardware potential. It does not describe the complete experience people will have after months of updates, mixed workloads and imperfect network conditions.

Xiaomi 18 Fold shown closed and open in red

Xiaomi 18 Fold shown closed and open. Image: Xiaomi / Weibo, as attributed by PTTL.

Begin with the boundary of the claim

“On-device AI” can mean several different things. A feature might run an entire model locally, use a small local model before sending a request to a server, cache data on the device, or simply rely on an AI-branded cloud service. Each case has different implications for latency, data handling and offline use.

Ask a concrete question for each feature: which model executes locally, which inputs leave the device, what account or network connection is required, and what happens when the connection disappears? A local summarization feature that uploads a document for final processing should not be described as fully local. Conversely, a function that works in airplane mode with no external account request offers stronger evidence of local execution.

Manufacturers should be able to explain these boundaries in product documentation. If they cannot, treat broad privacy language as a product claim awaiting verification, not as a completed privacy assessment.

Test a workflow instead of a benchmark

The XRING O3's announced performance is a reason to test real work, not a substitute for that work. Synthetic benchmark scores are controlled snapshots. They cannot show whether a phone stays responsive while a person reads reference material, switches between applications, records notes and runs a local feature for several minutes.

Choose one repeatable workflow. For example, open a long PDF or web page on the inner display, collect key points into a note, run the phone's available writing or translation function, then continue with a video call or navigation app. Repeat the sequence with the device cool, after sustained use and while charging. Observe delays, app reloads, heat, battery loss and whether the folded and unfolded layouts preserve the user's place.

The result is more useful than a single score because it connects hardware to an activity. It also produces a comparison that another reviewer can repeat on a competing phone.

Check sustained speed and thermal behavior

Foldable phones have a difficult packaging problem. They distribute a processor, batteries, cameras, radios and a hinge across thin halves. A device can deliver impressive peak performance and still reduce speed quickly when it becomes warm. That is not automatically a failure; thermal limits protect the device. It is a reason to measure sustained behavior rather than extrapolate from a short test.

Run the same AI task several times, then repeat it after a period of gaming, camera recording or multitasking. Record completion time, surface temperature, screen brightness changes and whether the system disables or slows a feature. Compare battery percentage before and after a fixed test interval rather than relying on an estimated remaining-time display.

The useful question is not whether a processor is fast in isolation. It is whether the phone can maintain a comfortable, predictable workflow at the temperature and battery level where people actually use it.

Treat a larger screen as an interaction test

A wide foldable can make side-by-side reading, note-taking and comparison less cramped. The benefit depends on application behavior. Check whether an app retains its scroll position when the phone opens, whether split-screen panes can be resized, and whether text fields, keyboards and menus stay reachable.

Also test the less glamorous states: receiving a call, changing orientation, switching from Wi-Fi to mobile data, locking and unlocking the phone, and opening a large attachment. A productivity feature has little value if it loses draft text or forces the user to restart after an ordinary transition.

The launch reports describe the 18 Fold as a China-first device. That matters for evaluation, too. Availability, local services, language support, carrier behavior and warranty policies can alter which AI features a buyer can actually use. Do not infer global software support from an announcement in one market.

Separate vendor claims from independent evidence

The historical context for Xiaomi's internal silicon is real: Xiaomi's own materials document the earlier XRING O1 launch in 2025. The new product reporting establishes that XRING O3 is part of the 18 Fold announcement. Neither source type is an independent long-term review.

Use clear labels when documenting a device. “Manufacturer states” is appropriate for memory bandwidth, model capability and theoretical throughput until an independent test verifies a relevant outcome. “Measured” should identify the test, software version, settings and conditions. “Unavailable” is an honest result when a promised feature cannot be accessed in the test region or account configuration.

This vocabulary prevents a common mistake: turning a launch claim into a conclusion about a buyer's experience. It also makes later updates easier because new measurements can replace a clearly bounded claim without rewriting the whole article.

Evaluate the support commitment

A local AI feature is not only a chip feature. It depends on operating-system updates, model revisions, drivers, application integrations and security maintenance. Ask how long the device receives platform and security updates, whether new on-device features will reach the model, and how a model or system update can be rolled back if it breaks a workflow.

For a foldable, repair and display durability belong in the same review. A large inner screen and a hinge can determine whether a device remains usable long enough for its software promises to matter. Early launch coverage cannot answer those questions; independent testing needs time.

The 18 Fold is a useful case study because it combines a new form factor, in-house silicon and local-AI messaging in one premium product. The right response is neither to dismiss the announcement nor to accept every specification as proof. Verify feature boundaries, test a realistic task under sustained load, inspect the folded and unfolded interaction, and record what remains unknown. That is how a promising hardware launch becomes evidence a buyer can use.

Editorial method

AI Tools Radar separates product facts, editorial judgment, and commercial placement. Updated facts retain their verification date.

Sources

Browse the directory