MiniLM-L6
Not run.
RAM DEVELOPMENT LAB
Second-pass Safari-only benchmark. It never reads RAM data. It uses only synthetic RAM-style text. Internet is used only to download test runtimes/models; inference runs on this iPhone.
Do not retry a model repeatedly if Safari reloads during it. Reopen the page; v3 records the last phase reached. Completed results survive reloads.
Not run.
All three encoders already passed the easy v2 screen. This uses 15 harder RAM-style retrieval probes with closer distractors and reports Top-1 accuracy, mean reciprocal rank, and average winner margin.
Not run.
Not run.
Not run.
Unlike v2, each model stays loaded but receives one Capture per inference, which matches how RAM would actually use a deep interpreter. Grounding is stricter: every extracted span must be an exact substring of the Capture, and temporal/condition fields must also be grounded.
Same model class that reached v2 output, but WebLLM is pinned to 0.2.82 and the harness is corrected.
Not run.
Very low stated WebLLM VRAM requirement (~711 MB); interesting quality-per-memory candidate.
Not run.
Stronger modern candidate. Pinned to WebLLM 0.2.82 for the stability comparison.
Not run.
Newer compact candidate; requires current WebLLM model support.
Not run.
~2.25 GB stated WebLLM VRAM. Run only if the earlier candidates complete without Safari instability.
Not run.
call mom tomorrow buy milk tonight finish report nowdon't buy milk I already got it but grab eggs tomorrowI'm exhausted but I still need to finish UWorld questions tonightafter Sarah sends the file review it then email Johnfind the glue gun and find my Apple Watchfinish 10 questions from set 6 then do set 8No results yet.