RAM Local Intelligence Spike v2
Isolated WebGPU/WebLLM feasibility harness. This is not RAM, not Build 13 R10, and cannot commit memories.
Starting harness…
Isolation rule: deploy this folder to a separate HTTPS origin from RAM. It contains no RAM database, MemoryService, Backup, lifecycle, provenance, or persistence code.
1 · Device capability
Harness identity
Build
WebLLM
Context tokens
Engine
WebLLM is intentionally pinned rather than tracking latest. Harness 2 prioritizes Qwen2.5 0.5B → Qwen3 0.6B → Qwen2.5 1.5B while keeping prior comparison models.
2 · Model runtime
No model loaded
3 · RAM atomic-meaning probe
Use Runtime smoke first. Then use Analyze plain JSON as the primary semantic test: it avoids grammar-guided decoding so a Qwen/xgrammar failure cannot masquerade as a model-quality failure. The constrained button is a separate runtime stress comparison.
Raw model output
No inference yet.
External validation
No inference yet.