RAM Local Intelligence Spike v2

Isolated WebGPU/WebLLM feasibility harness. This is not RAM, not Build 13 R10, and cannot commit memories.

Starting harness…
Isolation rule: deploy this folder to a separate HTTPS origin from RAM. It contains no RAM database, MemoryService, Backup, lifecycle, provenance, or persistence code.

1 · Device capability

Harness identity

Build
WebLLM
Context tokens
Engine

WebLLM is intentionally pinned rather than tracking latest. Harness 2 prioritizes Qwen2.5 0.5B → Qwen3 0.6B → Qwen2.5 1.5B while keeping prior comparison models.

2 · Model runtime

No model loaded

3 · RAM atomic-meaning probe

Use Runtime smoke first. Then use Analyze plain JSON as the primary semantic test: it avoids grammar-guided decoding so a Qwen/xgrammar failure cannot masquerade as a model-quality failure. The constrained button is a separate runtime stress comparison.

Raw model output

No inference yet.

External validation

No inference yet.

4 · Evidence

Runtime log