Testing and release evidence

Measure Time to First Token on Device

Define start and first-output events and keep cold and warm runs separate.

Last reviewed: 2026-09-16 · Fact IDs: BENCHMARK-01, THERMAL-01

Direct answer

Benchmark and regression content must be generated from versioned fixtures and raw run records, including failures and fallback.

This page does not publish benchmark or compatibility results. It shows the evidence required to answer time to first token on device without turning an assumption into a product claim.

Evidence to collect

The claim becomes reviewable only when the following evidence is attached to the same artifact and test run:

  • Clock boundaries
  • Cold and warm labels
  • All run records

A release claim is valid only for the recorded configuration and acceptance set. Generalization requires additional records.

Implementation workflow

  1. State the decision the test must support.
  2. Pin device, software, artifact, runtime, and backend.
  3. Version the workload and expected properties.
  4. Separate cold, representative, and sustained phases.
  5. Keep ordered raw records including failure.
  6. Publish method and evidence before conclusions.

Keep each transition observable. A failure should identify the stage, artifact, runtime, and recovery action without logging private user content.

Failure patterns to prevent

  • Starting after preprocessing
  • Publishing the fastest run

Also prevent silent fallback, unpinned artifacts, missing cancellation, and conclusions that combine unlike configurations. Store unsuccessful runs alongside successful ones.

Minimum reproducibility record

LayerRecord
DeviceManufacturer, model, chipset, RAM class, operating-system build
SoftwareApplication version and git commit
ModelFamily, variant, revision, format, file length, hash, quantization
RuntimeName, revision, requested backend, observed backend evidence
WorkloadFixture revision, input hash, prompt hash, output policy
OutcomeCompleted, failed, cancelled, fallback, and privacy-safe diagnostics

Release checklist

  • ☐ The primary query is answered without an unsupported number.
  • ☐ Every artifact and runtime is pinned.
  • ☐ The representative task and failure policy are explicit.
  • ☐ Lifecycle, cancellation, cleanup, and fallback are tested.
  • ☐ User-content and network boundaries are documented.
  • ☐ Result wording applies only to the recorded configuration.
  • ☐ The page links to raw method or evidence when results are added.

Sources and related evidence

Chinese deployment and troubleshooting content is organized in the 奇连 AI 端侧专题.