Stop paying for every token. Seriously. On-device LLMs deliver enterprise AI functionality with zero cloud costs and guarantee total user privacy by running entirely on the device.
The shift to on-device AI is real. Yes, cloud-based AI makes sense for training, but when it comes to using mobile devices for inference, the economics and privacy implications are starting to make it look way less attractive.
Of course, you won't be able to get the same performance from those on-device models just yet, but on-device GenAI is rapidly maturing. The demand for secure, fast AI-driven experiences is driving this. The future relies on continuous innovation in model compression and better collaboration between LLM researchers and mobile hardware developers. We need to overcome resource barriers and boost mobile adoption, and with everything that is happening in the industry over the recent years, that does seem achievable.
This isn't just theory though. We'll walk through three critical areas:
- privacy and cost efficiency,
- hardware limitations we're solving, and
- how RAG makes local LLMs actually useful.
Then we'll show you a working demo we built that summarizes text entirely on-device.
Overcoming mobile hardware limitations through optimization
The energy consumption problem
Deploying LLMs on mobile devices faces a fundamental challenge: substantial energy consumption that rapidly depletes battery life. Running high-demand scenarios like 20 iterations of large prompts, for example, can cause a 16% mean battery level delta on a typical device.
Mobile devices lack the thermal management systems for continuous, high-intensity computation. Your phone gets hot. Performance throttles. Users notice and complain, unfortunately.
Testing across various devices reveals a consistent pattern. Small models (under 1B parameters) are manageable, but scale up and you hit thermal and power limits quickly. CPU and GPU heat up, the device throttles to protect itself.
Model optimization through quantization
One technique to tackle the challenges related to using LLMs on mobile devices is model optimization through quantization, which is where the magic happens. It's the process of converting model weights from high-precision floating-point numbers (like FP32 or FP16) to lower-precision integers (like INT8 or INT4). FP32 uses 32-bit full precision with each parameter taking 4 bytes, FP16 uses 16-bit half precision taking 2 bytes, while lower-precision formats such as INT8 (1 byte) or INT4 (0.5 byte) reduce storage significantly. By reducing the precision of model parameters, we shrink model size dramatically. For example, a 7B parameter model that would typically occupy about 28 GB at FP32 precision drops to 14 GB at FP16, 7 GB at INT8, and just 3.5 GB at INT4. Going from FP32 to INT8 makes the model 4× smaller, while INT4 quantization achieves an 8× reduction. This massive size reduction enables LLMs to run on resource-constrained devices, with INT8 models running faster and using 4× less memory with only minor accuracy loss, and INT4 methods like QLoRA demonstrating approximately 99% of the original model's performance.
But size reduction is just part of it. The real win is computational efficiency.
INT8 quantization achieves a 40% reduction in computational cost and power consumption compared to FP16. INT4 quantization pushes that to 65% reduction. On edge devices, INT4 configurations show 4x throughput increases and cut power consumption by 60%.
These aren't theoretical numbers. Hardware evaluations on actual mobile devices confirm it. The trade-off is a slight quality degradation, but for most use cases, the precision loss is negligible: you're getting 95% of the quality with a fraction of the cost.
It's worth noting that quantization isn't just for small language models. True LLMs—models with billions of parameters—benefit significantly from quantization as well. The same size and efficiency gains apply, making even large models viable for on-device deployment where they would otherwise be impossible to run.
However, when it comes to quantization, choosing the right quantization strategy definitely matters. Text summarization works great with INT4. More complex reasoning tasks benefit from INT8. That's the sweet spot for most production deployments.
Thermal and performance considerations
Even with quantization, thermal management matters. Short bursts of inference work great. Continuous generation over minutes requires smart chunking and giving the device time to cool.
In the demo we built (more on that below), we process text in chunks and interleave small delays between operations. Battery life matters, but it's also about maintaining consistent performance without triggering thermal throttling. Once throttling kicks in, everything slows down.
The functional enhancement of retrieval-augmented generation (RAG) and integration
Improved output quality and relevance
RAG improves the quality of an LLM's responses by supplementing the model with relevant information retrieved from a local knowledge base. Accuracy matters, but so does relevance and context.
With local RAG, the model generates answers that are accurate, relevant, helpful, because it has access to domain-specific information. It doesn't need to rely solely on training data, which might be outdated or generic. This makes a real difference in production.
Structured generation workflow
The process is straightforward:
- Question: User query comes in
- Retrieve: Search the local knowledge base for relevant context
- Augment: Create a context-rich prompt combining the query and retrieved information
- Generate: LLM produces the context-aware final answer
What's interesting is how this maps to mobile constraints. The knowledge base can be pre-indexed and stored locally. Retrieval happens fast because it's just vector similarity search on embedded documents. The LLM then generates using that enriched context.
We're not talking about internet-scale search here. We're talking about a curated knowledge base that not only contains your company docs, product information, user manuals, but also fits on the device and gets queried in milliseconds.
Seamless on-device integration
The on-device LLM ecosystem is maturing with impressive speed. Libraries like react-native-executorch (which we use in this demo) enable LLM execution on-device using PyTorch's ExecuTorch runtime. Other options like react-native-ai provide compatibility with engines like MLC LLM and the Vercel AI SDK.
For iOS 18+, Apple Foundation Models provide Instant AI features natively. Text generation, embeddings, transcription, and speech synthesis—all on-device, all using system-level APIs.
The integration story is getting better. A year ago, getting an LLM to run on-device required custom native code and significant optimization work. Now you can pull in a library and have basic inference working in an afternoon.
The demo we built uses react-native-executorch with a quantized model. It's a React Native app that takes pasted text, fetches the content, and generates a summary entirely on-device. No API calls. No cloud infrastructure. No per-request costs.
Building the demo: text summarization
The demo application in this repository demonstrates on-device summarization. Users paste text directly, and the app processes it locally and generates a summary using an on-device LLM.

Architecture-wise, it runs the model entirely on-device, using optimized quantized weights. Inference stays fast and power-efficient.
Key implementation details:
- Text chunking for long documents (respects context windows)
- Local vector embeddings for RAG (if you want to add knowledge base features later)
- Progressive summarization for very long texts

The app allows switching between different models (see the model registry in the repo), and our testing revealed some interesting trade-offs. The smallest model, SMOLLM2_1_135M_QUANTIZED at just 50MB, runs incredibly fast but doesn't produce the best results. On the other end, going for the biggest model like LLAMA3_2_3B_SPINQUANT quickly leads to memory issues on many devices. From our experience, QWEN2_5_0_5B_QUANTIZED (~200MB) produces the most balanced result — good quality output with reasonable performance and memory usage.

You can try it yourself: clone the repo, install dependencies, run npm run ios or npm run android. Works on iOS and Android. Everything happens on the device.
Conclusion
The future of mobile intelligence may very well lie in the on-device first philosophy since even advanced features like Retrieval-Augmented Generation (RAG) can be run locally, ensuring privacy and resulting in zero cost by avoiding cloud server reliance.
Running Large Language Models (LLMs) on mobile devices may be challenging, especially with substantial energy consumption and battery life issues being real concerns. However, developers manage to find ways to overcome this. With optimization techniques like quantization, we can shrink model size by up to 68% and cut computational costs by up to 65%, which in its turn means that powerful, complex AI systems are now viable on average smartphones.
Hardware is getting better. Tooling is maturing. Economics are shifting in favor of local execution.
We're not at the point where every mobile app needs its own LLM. But for use cases where privacy, cost, or offline capability matter, on-device AI is already the right choice. The demo in this repo shows it's not just possible, it's practical.
The path forward requires continuous innovation in model compression, better collaboration between LLM researchers and mobile hardware developers, and tools that make on-device AI as easy to deploy as cloud-based alternatives.
But wait - there's more.
Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.
Insights
Perspectives on AI in engineering, product development, and strategy, for enterprise executives.
Community
Deep dives and tutorials by engineers, for engineers.
You may also like


