Built by Kato / PMZFX
I'm Kato Papairo. This guide brings together B70 benchmark work, CUDA-to-Intel investigations, and contributions to llama.cpp's SYCL backend. The goal is to make the next person's setup or diagnosis easier, with enough detail to challenge a result.
This is an independent project, not an Intel publication. Follow the PMZFX GitHub profile for public work, or start with the upstream contribution record. For technical collaboration, a focused issue or discussion on the relevant public repository is the most useful starting point.
Measurement standards
Every published benchmark retains its original test date, source identifier, source-file hash, software configuration, and limitations. Editorial review is tracked separately. Reviewing an April result in September does not make it a September benchmark.
The current dataset contains five April records and one June configuration. Original values and reported standard deviations are preserved in the JSON and CSV exports. The UI rounds rates to two decimal places. Unknown metadata is null in JSON and blank in CSV; it is never replaced with a plausible default.
- Prompt processing: the explorer consistently shows a 512-token prefill workload.
- Decode: a 128-token generation microbenchmark. June's raw row explicitly uses zero prompt tokens and zero depth.
- Context: configured capacity and exercised prompt length are separate fields.
- Reproduction: April's build is dirty and the exact local diff is unavailable. Thread count follows each result's metadata.
- Environment: cohort-level software descriptions are labeled as such, rather than presented as per-run captures.
Downloads include source attribution and caveats alongside the measurements. Public build inputs are limited to selected site content and curated data. The private research archive is not served.
What is not a ranking
Different models, quantizations, GPU counts, cache settings, and backend versions answer different questions. A throughput bar is not a model-quality comparison, and April-to-June differences are not an isolated software speedup.
Energy rankings are withheld. Some archived summaries include loading and idle power from another card; dividing decode rate by those averages does not establish decode-only efficiency. GPU telemetry is also not whole-system wall power. Zero VRAM telemetry is treated as missing data.
The preserved hardware prose disagrees about negotiated PCIe generation. Current prices, regional availability, and price/performance are not established by the old April price snapshot.
Review queue
- Installation: produce a current version-pinned baseline and validate real inference, repeated requests, and long-input behavior.
- April evidence: recover the dirty diff and reconcile thread, weight-size, and topology discrepancies.
- June evidence: recover missing OS/kernel/driver details and build cleanliness.
- Compatibility: publish an independently dated alpha checkpoint with public harness evidence before expanding support claims.
- Media: replace the clearly labeled schematic with an authorized original card photograph when available.
- Video: recover per-run timing and output evidence before publishing Wan performance comparisons.
Historical means a dated observation. Experimental means the scope and stability remain limited. Source reviewed means current documentation or external reports were checked, with no local execution implied. Needs retest means current execution has not been validated. None is a substitute for reading the environment and limitations.
Corrections improve the guide
Include a link to the claim, the configuration it applies to, and the smallest supporting reproduction or source. Open a correction in the benchmark repository ↗
Reviewed September 10, 2026. No GPU benchmark or service change was performed as part of building this website.