← All blog posts

Porting a typed-decision engine to Apple's foundation models: a thousand tests and the limits

September 23, 2026 · AI

A support ticket reading "I was billed twice this month. Please refund the duplicate charge." goes to Apple's on-device model with three typed questions. The framework exposes no logits, so the model is asked to write an answer first and then a weight from 0 to 100 for each of four departments. What comes back is billing at 1.000 and the other three at 0.000, with urgency and the refund question also at 1.000. Below, three panels. What the live tests found: asked for the weights alone, German, Spanish and Italian all came back french, the second option, and answering first took accuracy from 80.6% to 91.2%. What the platform costs: about 410 ms before the first token, about 90 tokens a second, one session kept in cache, 8,192 tokens of context, and Private Cloud Compute refused from a command-line tool with error 1046. Against Open-Jev, on the customer-service workload: jev-mac 5,506 ms, GPT-6 Astra 1,938, GPT-5.6 Luna 918, Jev 295 and Open-Jev 2B 85; and 172 of 300 FizzBuzz decisions against Jev's 299.
The model cannot hand over its probabilities, so jev-mac asks it to write them down. For the README's first example it writes 1.000, and it happens to be right. Underneath, what a thousand tests and Open-Jev's benchmarks had to say about everything else.

A few minutes after midnight on 23 September, I pasted a link into Claude Code and asked for one thing; the architecture of laya-mlx, built on Apple's foundation models and nothing else. By half past eight in the evening the repository was public on GitHub, with an engine, a thousand tests, a snake game and a README full of measurements (a few of them flattering, most of them not).

I did not write a line of it. Claude Code wrote the engine and the command-line tool (about 3,100 lines of Swift), the tests (2,700 more) and the README. I asked questions, played the snake and read the numbers. A few of those questions changed the project, and they are a good part of this post. The rest is about Apple's models, and what they will and will not do for you.

What laya does, and what Apple gives you

laya-mlx is an independent port of Laya, a family of typed-decision models from Convai Innovations, to MLX on Apple silicon. You give it a state (a support ticket, an email, a snake board) and a set of typed questions. A choice returns probabilities over named options, a score returns probabilities over the levels of a rubric, and a noul returns the probability that a proposition is true. Underneath there is an encoder with decision heads; one forward pass, and the probabilities come straight off the heads without generating a single token. Its README reports 13.4 ms for a short English decision on an M3 Max. (I have not run it myself, so that number is theirs, not mine.)

Apple's FoundationModels framework offers something else. You open a session with a language model that lives inside macOS, you send it a prompt and you get text back. Or, and this is the useful part, you get text shaped by a schema; guided generation constrains the decoding, so the answer always parses into the structure you asked for. What you never get is a probability. The framework's interface, as the macOS 27 SDK ships it, has no logits and no log-probabilities (search it for logprob, nothing comes back), and the only probability in it is the threshold for nucleus sampling.

So jev-mac asks the model to write its probabilities down. For every question it builds a schema that wants the best option first, as an enum, and then an integer weight from 0 to 100 for every option. The weights are normalised into a distribution, a temperature is applied per question, and the confidence is one minus the normalised entropy, as in laya. There is a second head, vote, which samples the answer several times and counts. It is slower, but its numbers come from sampling rather than from what the model says about itself.

For what it is worth, jev-mac is in good company. Open-Jev, whose benchmarks I checked it against later, has the same problem with the GPT models it measures; its page describes their columns as "verbalized probability vectors in constrained JSON, not token logprobs". Same trick, same caveat.

And there is something here that I found oddly familiar. I spent my PhD on J%, which turns an SQL query or a regular expression from a string that fails at runtime into a type the compiler checks. Guided generation does the same to a model's answer. The schema guarantees the shape; every probability sits between 0 and 1, they sum to 1, the decision is the argmax, and in every run of the test suite that held, with zero invalid outputs. What a type never told you, in 2015 or now, is whether the value inside it is right.

Here is the first example in the README, run again today:

$ jev-mac predict --preset triage --summary "I was billed twice this month. Please refund the duplicate charge."
department [choice] → billing  confidence=1.00
    billing    ████████████████████████  1.000
    technical  ························  0.000
    sales      ························  0.000
    account    ························  0.000
urgency [score] → level 1  expected=1.00  confidence=1.00
    1  ████████████████████████  1.000
    2  ························  0.000
    3  ························  0.000
    4  ························  0.000
wants_refund [noul] → true  P(true)=1.000  ████████████████████████

on-device · 3 calls · 1596 in / 0 cached / 115 out tokens · 2961 ms

Right on all three questions, and certain on all three. Keep that certainty in mind; it comes back.

A thousand tests, and a slot called french

Less than half an hour after the first request, I asked for a thousand tests. Claude Code split them into two layers. 577 are deterministic; they drive the JSON parser, the prompt builder, the read-out and the cache with inputs shaped like the model's, run in under a second and never call the model. The other 423 are live and take about eight minutes against the on-device model; 306 labelled states (every preset, plus language identification, reading comprehension and numbers inside JSON), 90 snake positions and 27 end-to-end checks. (The test bundle also refused to be code-signed at first, because iCloud Drive decorates the folder with extended attributes. The tests are built in /tmp now. Nobody warns you about that one.)

The live layer paid for itself on its first run. Asked for the weights alone, the model piled the mass onto an early option whatever the text said. German, Spanish and Italian sentences all came back french, the second option in the list, and when the options were reversed, the answers moved with the slot. The model was answering the list, and the text hardly mattered.

That is why the schema asks for the answer first. Committing to an option before writing a single number lifted accuracy on choice questions by 9 points and on score questions by 14, and it cost about 29% in latency (a three-question prediction went from 1.87 s to 2.43 s). Asking the noul questions as yes or no, and rewording one snake question, halved the Brier score of the noul probabilities. The snake needed one more fix, and that one deserves its own section.

Table 1. The live suite after each change to the prompts; 396 labelled cases on the on-device model, Apple M4 Max, macOS 27.0. Accuracy in percent, the Brier score of the noul probabilities (lower is better) and the median latency per question in milliseconds. There were no invalid outputs in any row. The best figure in each column is in bold; the fastest engine was also the least accurate.
EngineAccuracy (%)noul BrierMedian per question (ms)
Weights only80.60.202732
+ the answer first86.10.208937
+ yes or no for noul, one snake question reworded90.90.099939
+ the snake state in plain words91.20.093930

91.2% on 396 labelled cases, and not one invalid output in any row. I would like to stop there, but the README does not, and neither will I. The misses guided the prompt changes, so the last row is optimistic. Six of the eight scam and phishing emails get a spam probability of exactly 0. And many of the misses come back at P = 1.00; the model is often as certain when it is wrong as when it is right, which is exactly the property a decision engine should not have, because nothing downstream can tell the two apart.

Are you sure you did not reload the model?

Around noon I asked the question in the heading. Every question was taking the best part of a second, and on a local model that looks a lot like something being loaded and thrown away on every call.

It was not. The model lives in a system service. modelmanagerd owns it, and it runs in a process of its own (TGOnDeviceInferenceProviderService, if you want to look for it) that used 1.0 to 1.7 GB during our runs, kept one PID for more than eight days and logged no load or unload. The model stays resident between calls, whoever makes them.

What jev-mac did do was open a fresh session for every call, which throws away any chance of the service reusing the prompt. So it switched to pooling; each question keeps its own sessions, and a session serves one request, is reset to its instructions and waits for the next request with the same question. The cache started to hit. The whole thing got between 1 and 3% faster.

Why so little? Two measurements explain it. Every request spends about 410 ms before its first token, and it spends the same with 0 or 220 cached tokens, and with 239 or 437 input tokens; reading the prompt is simply not where the time goes. And the service keeps the context of only the most recently used session. Alternating two sessions gave zero cached tokens on every call, and the same session twice in a row gave 220 on the second. A prediction runs its questions interleaved, so they evict each other. (Anyone who wrote Java in the 2000s has done this at least once; pool everything, measure, and find out that the pool was never the problem.)

The rest is generation. An answer with its weights is about 41 tokens, at roughly 90 tokens a second, so around 450 ms. One question costs about 0.9 s, the three-question triage about 2.4 s. Laya's figure for one question, remember, was 13.4 ms.

The snake that played at random

laya-mlx ships a terminal snake, so jev-mac got one too. At every step the engine works out the facts of each move (is it legal, does it leave room for the body, does it eat the food or get closer to it), the model chooses, and a safety shield executes the most probable move that is legal and leaves room for the body. The tests passed.

Then I played it, and wrote back: "have you tested snake, it does not work, it is like it is random".

It had been tested. The tests checked that every move was safe, and every move was; nothing checked that the snake was going anywhere. The model was handed raw numbers for each move (food_distance_after: 6, free_space_after: 188 and so on), and it cannot compare numbers across options. So it fell back on the first option, UP, and circled a corner without ever reaching the food.

The fix was to stop giving it numbers. The state is now written in words, a line per move:

Snake on a 16×12 board, length 4, heading RIGHT. Head at (5,4), food at (9,4), 4 steps away.
Moves:
- UP: safe, moves away from the food (5 steps)
- DOWN: safe, moves away from the food (5 steps)
- LEFT: not possible (the snake's own body)
- RIGHT: safe, gets closer to the food (3 steps)
The food can be reached after moving UP, DOWN, RIGHT.

Over the same 200-move game (seed 7, one question per move), a random safe move eats 1 food, the numbers version ate none in its first 40 moves, and the words version eats 18. It heads for the food on 194 of the 194 moves where that was safe, and the live test now checks exactly that, not just that the move is safe. In the applets post last week I wrote that finding an oracle is the step that pays. Here the oracle was me, playing the game.

A terminal window running jev-mac snake. On the left, a 16 by 12 board where a green snake with a yellow head moves towards a red square of food, eats it and grows. On the right, the game's panel, titled JEV-MAC SNAKE, AFM 3 Core Advanced on Apple M4 Max: the model's probabilities for the next move, up, down, left and right, drawn as bars with the chosen move marked, then the dead-end risk and whether the food is reachable, and a decision time of about 2.4 seconds a move.
jev-mac snake with all three questions, recorded on this Mac on 23 September 2026: 45 moves, four foods, no deaths, and a step towards the food on every move where that was safe. It plays back at 0.35 s a move; the real decisions took about 2.4 s each, which is why the game's own panel reports 0.41 moves a second.

So does the model play snake now? In a way. A greedy rule that reads the same facts, and takes the safe move closest to the food, eats 17. The words already contain the answer; the model reads our analysis back to us and finds one more food. It also takes about a second per move with one question, and 2.4 s with all three. The laya-mlx README reports about 75 moves a second for its snake, on an M3 Max.

Against the published numbers

I also asked Claude Code to check the benchmarks that Open-Jev publishes. Open-Jev measures its own open checkpoints, the hosted Jev service and two GPT models on typed-decision workloads. Most of its suites use inputs that are not public, so only two could be rebuilt, and only in shape; a latency matrix, and FizzBuzz as a control.

Table 2. jev-mac against the figures Open-Jev publishes (read on 23 September 2026). The two latency columns are the median of 20 timed requests after 3 warm-ups at concurrency 1, in milliseconds: the customer-service workload asks 8 questions, the other puts 32 candidates over a 1,024-token state. jev-mac runs on an M4 Max, Open-Jev 2B on a local H100, and the hosted services include the network. FizzBuzz counts correct answers out of 300 typed decisions. jev-mac's workloads are rebuilt in the same shape as Open-Jev's, not from the same texts, so read the table as orders of magnitude. The best figure in each column is in bold.
SystemCustomer service (ms)32 candidates (ms)FizzBuzz
jev-mac5,5066,681172
Open-Jev 2B851,016pending
Jev 1.13.0295301299
GPT-5.6 Luna918690300
GPT-6 Astra1,9381,388300

On the customer-service workload jev-mac is 65 times slower than Open-Jev's 2B checkpoint on an H100, 19 times slower than the hosted Jev and about three times slower than GPT-6 Astra; and the hosted services pay for the internet in their figures, while jev-mac pays for nothing but the model. The candidate count is what hurts. A weight is written out for every option, so 32 candidates means about 290 output tokens and six seconds, whatever the size of the state.

FizzBuzz is the result I keep coming back to. Three typed questions for every integer from 1 to 100; divisible by 3, divisible by 5, and what FizzBuzz prints. jev-mac gets 172 of the 300 decisions right (57.3%), against 299 for Jev and 300 for both GPT models. Two of the three questions score below the majority-class baseline; answering "no" to every divisible-by-5 question would get 80 of 100, and the model gets 57. The on-device model cannot do this arithmetic reliably, and no schema on the output changes that.

This is also where I asked for one more thing. A number without its machine is a rumour, so every benchmark now opens by printing the machine and the conditions of the run, repeats the conditions at the end, and warns when the machine is busy, on battery or throttled:

machine     Apple M4 Max (Mac16,5) · 16 CPU cores (12 performance + 4 efficiency) · 40 GPU cores · 64 GB memory
conditions  macOS 27.0 (26A428) · AC power (battery 80%) · Low Power Mode off · thermal nominal · load 2.14 · memory pressure normal
model       AFM 3 Core Advanced (on-device) · 8,192-token context

Where Apple's models stop

To be fair to Apple, nobody promised probabilities. The framework is built for the features Apple ships on top of it (summaries, tagging, the writing tools) and for apps that want structured output, and a decision engine is a different job. But the list of what stood in the way is longer than I expected, and it splits in two; what the model cannot do, and what the platform will not let you do.

The model first.

  • No probabilities. They are written, not read, and they come out certain. The triage example above is 1.000 across the board, and so are many of the misses.
  • A small model. It cannot do FizzBuzz arithmetic, and it cannot compare numbers across options, which is why the snake needed words.
  • 8,192 tokens. Instructions, prompt, schema and answer all share them.
  • Fifteen languages. 24 locales in all, and Greek is not one of them. jev-mac marks a Greek state as unsupported, so I cannot ask it anything in my own language.
  • No reasoning on the device. The capability flag says so, plainly.

Then the platform.

  • A floor under every request. About 410 ms before the first token, one request per question, and every weight written out at about 90 tokens a second.
  • A cache that remembers one session. Hence the 1 to 3%.
  • Guardrails you cannot relax for a typed answer. The permissive setting only covers answers generated as plain text; the SDK's own documentation says that anything else behaves as in the default mode. The defaults blocked 5 of the 20 moderation inputs in the suite (harassment, self-harm, adult content), which is exactly what a moderation preset exists to classify.
  • No fine-tuning. Custom adapters were deprecated in macOS 26.4 and are obsoleted in 27.0, according to the SDK. Laya's heads were trained for this job; here you have prompts, and nothing else.
  • No Private Cloud Compute from the command line. It reports itself available, with a 32,768-token context and reasoning, and then rejects every request from an unsigned binary with ModelManagerError 1046. A signed app with the right entitlement would probably get through; I have not tried.
  • A model you do not own. It ships with the operating system, it changes when the operating system changes, and you cannot pin it. Every jev-mac figure in this post belongs to macOS 27.0 (26A428), on one M4 Max.

Epilogue

The bill is the part I keep thinking about. Last week I had two Java applets from 1999 rebuilt, because the runtime they needed was a browser plugin that browsers one day stopped running. jev-mac has the same shape of problem, in a newer box. laya-mlx keeps its weights on Hugging Face; the weights are a file, and you can keep a copy. jev-mac keeps nothing. Its model belongs to Apple and is updated with macOS, and the day it changes, every figure in the README changes with it, without a commit anywhere. That is why the benchmarks print the build now, and it is why the first line of the README says experimental.

The code is on GitHub as bkarak/jev-mac, under the MIT licence, for fun and research.