Our new on-device model looked too big for a tight iOS extension. It wasn't. How a hand-written decision memo, not more code, led us to a safer fix.

Last week, I wrote an internal message explaining why we appeared to have only two bad options.
We were trying to run a newer neural model inside an unusually memory-constrained iOS extension. Our previous tree-based models had fitted comfortably. The new architecture, designed to support richer on-device contextual intelligence, was roughly an order of magnitude larger.
The obvious conclusion was that the model had become too large for the environment.
That conclusion was wrong.
The model could run within the available memory. What exceeded the limit was a transient optimization step performed while preparing it for execution.
Less than an hour after I wrote the message, we had an end-to-end implementation running inside the extension. The path from “two bad options” to a working third one taught us something about on-device AI - and something equally useful about coding with AI.
When engineers discuss whether an on-device model will fit, we tend to focus on the artifact and the inference workload:
Those are necessary questions. They were not the decisive questions in this case.
Our measurements showed a temporary allocation of several megabytes when the already compiled Core ML model was first loaded and optimized. In an ordinary application process, that amount would barely be noticeable. Under a single-digit-megabyte ceiling, it was enough for the operating system to terminate the extension before inference began.
This distinction matters. Reducing the model would not necessarily solve the problem, because part of the spike appeared largely independent of the model’s size.
The deployment lifecycle - not merely the deployed model - had become the bottleneck.
That left us with two apparent options.
The first was to preserve a separate generation of compact, tree-based models for this environment. That was the safer engineering choice, but it would mean maintaining an additional training pipeline, deployment path and set of interfaces. It could also prevent this integration from benefiting from our newer architecture, although the size of that performance trade-off had not yet been measured.
The second option was more inventive and much less comfortable.
During debugging, we discovered that the platform placed an optimized model artifact in an internal location after the main application opened the model. In a local test, we could copy that artifact into the extension’s accessible storage and run inference with low memory consumption.
It worked. It was also the kind of solution that should make an engineer nervous.
The location was undocumented. We had tested it on one device and one operating-system version. We did not know whether it would survive an upgrade. A failed fallback could put the extension into a crash loop. Even without calling a private API, depending on an internal implementation detail was difficult to defend as a durable SDK design.
A passing test is evidence that something can work. It is not evidence that it is safe to ship.
An AI coding assistant had helped produce the initial Core ML implementation and investigate the failing tests. In the same working context, it helped discover the internal optimized artifact and showed that copying it could get the model running.
That was useful work. But it also shaped everything that followed.
Once the session had accumulated the implementation, test results and cache-copying experiment, further questions about alternatives remained close to that solution. The conversation became: How can we make this copying mechanism safer?
That is a reasonable question. It was not the best question.
I stopped coding and wrote the problem out for the team. The message described the memory constraint, our observations, both proposed paths, the maintenance implications and the risks I could see. I wrote it by hand. For me, it was an exercise in rubber duck debugging: writing the message forced me to think through the problem and verbalize trade-offs that, until then, had existed only intuitively in my head. By the time the message was finished, the problem itself was better defined.
While waiting for a colleague to respond, I pasted that memo into a fresh session using the same AI model and asked for alternatives.
Without the accumulated momentum of the earlier debugging session, it challenged the internal-cache approach and suggested controlling model compilation directly through BNNS Graph.
That changed the problem.
BNNS Graph is part of Apple’s Accelerate framework. It can compile and execute graphs derived from Core ML model files and is intended for CPU-based inference where applications need tighter control over execution and memory allocation. Apple explicitly describes it as an option for latency-sensitive inference with strict memory-management requirements. Apple introduced the API in detail at WWDC24.
The most important capability for our case was control over the compilation step. We could perform that work in the main app, then copy only the final compiled and optimized graph to the extension. This kept the memory-intensive preparation out of the constrained process while avoiding any dependency on an undocumented system cache. Apple documents the relevant BNNS Graph compilation controls here.
This provided an official primitive for separating preparation from constrained execution. Instead of depending on a private cache created as a side effect, our implementation could explicitly control compilation and the resulting artifact.
We brought that finding back into the original coding session, tested it and then used a fresh session to implement the design from a written specification. In our initial end-to-end test, the new model executed inside the extension without the previous memory spike.
That is not yet a universal result. It is one implementation tested within a defined environment. We still need a broader device and operating-system matrix, failure-path testing and production evidence.
But the nature of the solution changed. We had moved from undocumented platform behavior to a public API designed for controlled memory use.
It would be easy to tell this as a story in which an AI assistant solved the problem in minutes. That would miss the useful part.
AI contributed substantially:
But AI also followed the shape of the context we gave it. The first session had evidence that an internal cache could be copied, so it kept improving that path. The fresh session received a decision problem rather than a coding history. That made it more likely to question the premise.
The human contribution was not simply approving generated code. It was recognizing that the current direction was unsafe, making the trade-offs explicit, changing the information environment and testing the proposed alternative on a real device.
The decision memo turned out to be the most productive prompt in the process.
The broader lesson is that deploying on-device contextual intelligence is not only a machine-learning problem.
A model can perform well in offline evaluation and still be unusable in an application because of process limits, initialization behavior, energy constraints or platform lifecycle rules. Those limits are not peripheral. They shape which intelligence can reach the moment when an application needs to decide whether to act, wait or get out of the way.
For a situational-awareness layer between smartphone sensors and application decisions, “runs on the device” is too weak a definition of success. The system has to coexist with the application, respect the operating environment and remain dependable under the constraints of the least forgiving process in the architecture.
There is also a lesson here for AI-assisted engineering. Long coding contexts are powerful because they preserve detail. They can also preserve assumptions.
When an AI session keeps making a risky approach more sophisticated, the next prompt may not need more logs or more code. It may need a concise account of what you believe, what you observed, which alternatives you see and what would make each one unacceptable.
Sometimes the best way to make progress with a coding model is to stop coding and write the decision down.