NEW! 50% revenue boost for Blackbox - read the case study!
Engineering

The Prompt That Solved Our On-Device AI Memory Problem Wasn’t Code

Our new on-device model looked too big for a tight iOS extension. It wasn't. How a hand-written decision memo, not more code, led us to a safer fix.

Reinhard Hafenscher

September 29, 2026

Last week, I wrote an internal message explaining why we appeared to have only two bad options.

We were trying to run a newer neural model inside an unusually memory-constrained iOS extension. Our previous tree-based models had fitted comfortably. The new architecture, designed to support richer on-device contextual intelligence, was roughly an order of magnitude larger.

The obvious conclusion was that the model had become too large for the environment.

That conclusion was wrong.

The model could run within the available memory. What exceeded the limit was a transient optimization step performed while preparing it for execution.

Less than an hour after I wrote the message, we had an end-to-end implementation running inside the extension. The path from “two bad options” to a working third one taught us something about on-device AI - and something equally useful about coding with AI.

The hidden cost was not inference

When engineers discuss whether an on-device model will fit, we tend to focus on the artifact and the inference workload:

  • How large is the model on disk?
  • How much memory do its weights occupy?
  • What are the input and output tensor sizes?
  • How much working memory does inference require?

Those are necessary questions. They were not the decisive questions in this case.

Our measurements showed a temporary allocation of several megabytes when the already compiled Core ML model was first loaded and optimized. In an ordinary application process, that amount would barely be noticeable. Under a single-digit-megabyte ceiling, it was enough for the operating system to terminate the extension before inference began.

This distinction matters. Reducing the model would not necessarily solve the problem, because part of the spike appeared largely independent of the model’s size.

The deployment lifecycle - not merely the deployed model - had become the bottleneck.

That left us with two apparent options.

The first was to preserve a separate generation of compact, tree-based models for this environment. That was the safer engineering choice, but it would mean maintaining an additional training pipeline, deployment path and set of interfaces. It could also prevent this integration from benefiting from our newer architecture, although the size of that performance trade-off had not yet been measured.

The second option was more inventive and much less comfortable.

During debugging, we discovered that the platform placed an optimized model artifact in an internal location after the main application opened the model. In a local test, we could copy that artifact into the extension’s accessible storage and run inference with low memory consumption.

It worked. It was also the kind of solution that should make an engineer nervous.

The location was undocumented. We had tested it on one device and one operating-system version. We did not know whether it would survive an upgrade. A failed fallback could put the extension into a crash loop. Even without calling a private API, depending on an internal implementation detail was difficult to defend as a durable SDK design.

A passing test is evidence that something can work. It is not evidence that it is safe to ship.

The first AI session found a local optimum

An AI coding assistant had helped produce the initial Core ML implementation and investigate the failing tests. In the same working context, it helped discover the internal optimized artifact and showed that copying it could get the model running.

That was useful work. But it also shaped everything that followed.

Once the session had accumulated the implementation, test results and cache-copying experiment, further questions about alternatives remained close to that solution. The conversation became: How can we make this copying mechanism safer?

That is a reasonable question. It was not the best question.

I stopped coding and wrote the problem out for the team. The message described the memory constraint, our observations, both proposed paths, the maintenance implications and the risks I could see. I wrote it by hand. For me, it was an exercise in rubber duck debugging: writing the message forced me to think through the problem and verbalize trade-offs that, until then, had existed only intuitively in my head. By the time the message was finished, the problem itself was better defined.

While waiting for a colleague to respond, I pasted that memo into a fresh session using the same AI model and asked for alternatives.

Without the accumulated momentum of the earlier debugging session, it challenged the internal-cache approach and suggested controlling model compilation directly through BNNS Graph.

That changed the problem.

Moving the expensive work to the right environment

BNNS Graph is part of Apple’s Accelerate framework. It can compile and execute graphs derived from Core ML model files and is intended for CPU-based inference where applications need tighter control over execution and memory allocation. Apple explicitly describes it as an option for latency-sensitive inference with strict memory-management requirements. Apple introduced the API in detail at WWDC24.

The most important capability for our case was control over the compilation step. We could perform that work in the main app, then copy only the final compiled and optimized graph to the extension. This kept the memory-intensive preparation out of the constrained process while avoiding any dependency on an undocumented system cache. Apple documents the relevant BNNS Graph compilation controls here.

This provided an official primitive for separating preparation from constrained execution. Instead of depending on a private cache created as a side effect, our implementation could explicitly control compilation and the resulting artifact.

We brought that finding back into the original coding session, tested it and then used a fresh session to implement the design from a written specification. In our initial end-to-end test, the new model executed inside the extension without the previous memory spike.

That is not yet a universal result. It is one implementation tested within a defined environment. We still need a broader device and operating-system matrix, failure-path testing and production evidence.

But the nature of the solution changed. We had moved from undocumented platform behavior to a public API designed for controlled memory use.

What the AI contributed - and what it did not

It would be easy to tell this as a story in which an AI assistant solved the problem in minutes. That would miss the useful part.

AI contributed substantially:

  • It wrote the initial integration.
  • It accelerated experiments around the failure.
  • It helped identify where the memory spike occurred.
  • A fresh session surfaced an API we had not considered.
  • It converted the resulting design into an implementation.

But AI also followed the shape of the context we gave it. The first session had evidence that an internal cache could be copied, so it kept improving that path. The fresh session received a decision problem rather than a coding history. That made it more likely to question the premise.

The human contribution was not simply approving generated code. It was recognizing that the current direction was unsafe, making the trade-offs explicit, changing the information environment and testing the proposed alternative on a real device.

The decision memo turned out to be the most productive prompt in the process.

On-device AI is integration engineering

The broader lesson is that deploying on-device contextual intelligence is not only a machine-learning problem.

A model can perform well in offline evaluation and still be unusable in an application because of process limits, initialization behavior, energy constraints or platform lifecycle rules. Those limits are not peripheral. They shape which intelligence can reach the moment when an application needs to decide whether to act, wait or get out of the way.

For a situational-awareness layer between smartphone sensors and application decisions, “runs on the device” is too weak a definition of success. The system has to coexist with the application, respect the operating environment and remain dependable under the constraints of the least forgiving process in the architecture.

There is also a lesson here for AI-assisted engineering. Long coding contexts are powerful because they preserve detail. They can also preserve assumptions.

When an AI session keeps making a risky approach more sophisticated, the next prompt may not need more logs or more code. It may need a concise account of what you believe, what you observed, which alternatives you see and what would make each one unacceptable.

Sometimes the best way to make progress with a coding model is to stop coding and write the decision down.

‍

You might find these useful as well

Engineering

The Prompt That Solved Our On-Device AI Memory Problem Wasn’t Code

September 29, 2026

Engineering

OpenClaw <> ContextSDK: AI Agents That Know You’re Walking

February 23, 2026

Engineering

Timing Your Purchasely Paywalls with ContextSDK: The Missing “When” in Paywall Targeting

February 18, 2026

Engineering

Show Paywalls at the Perfect Moment with ContextSDK and Superwall

November 14, 2025