Skip to main content

Portable quality

Portable quality is an unreleased candidate workflow for evaluating work that does not begin in a Git repository, such as documents, spreadsheets, exported conversations, and structured data. It is currently Planned: the source and local verification described here are not a production-availability claim.

Candidate workflow

The candidate /quality workspace and shared API contract can:
  • create a quality context and immutable revisions;
  • upload resources in resumable chunks, ingest supported files, and query the normalized rows;
  • submit checks against a pinned revision and evidence manifest;
  • propose and review reusable standards;
  • create scoped grants and controlled delivery records; and
  • enforce retention, revocation, deletion, and audit lifecycle controls.
The generated TypeScript and Python clients and their CLIs expose the same shared contract. Candidate MCP tools expose bounded reads and mutations under exact scopes. These clients and tools have local clean-install and protocol proof, but have not been claimed as published registry packages or deployed production interfaces.

How checks decide

Each requirement is decided by the evaluator it names:
  • Text assertions run locally. A match found in extracted text is proof. An absent match is proof only when the parser recorded exact coverage; PDFs, images, HTML with scripts or visual elements, and Office files with headers, drawings, charts, or formatted values are recorded as partial.
  • Calibrated LLM judges run a stored judge config. A config qualifies for a task family only through a calibration run of that exact config version, against a calibration set for the same task family, with human agreement of at least 0.8 and at most 5% unparseable judgements. The plan pins the config hash; an edited or deleted config is never substituted. On partial text a verdict is decisive only when the requirement’s evidence polarity allows it: a presence requirement can pass and an absence requirement can fail; a holistic one needs complete text. Mandatory controls cannot use a judge.
When automated evidence cannot decide a requirement, the check finishes waiting for evidence and lists each pending requirement with its reason. Attesting those requirements (POST /api/portable-quality/checks/{id}/evidence) re-queues the check through the normal worker: decisive automated evidence still wins, prior judge verdicts are reused rather than re-billed, and the attestation is recorded on the finding. The check’s principal may attest; a mandatory control needs an organization admin other than the principal.

Privacy and execution boundaries

Resource ingestion and query enforce the resource’s current processor and destination policy before reading stored content. Provider-bearing execution also intersects the current organization provider policy, every current resource destination policy, the frozen plan policy, and the current grant. A changed or missing permission fails closed before provider execution, and rejected telemetry records stable codes and identifiers rather than resource payloads. The candidate includes a local parser and exact structured query engine. Parser and query workers run in an empty network namespace where the host allows it; see docs/security/portable-sandbox-network-isolation.md. Judge calls go through EvalGate’s judge service and model gateway, so centralized provider-policy and egress controls still apply. Optical character recognition, embeddings, and live external-provider execution against a named provider are not currently proven.

Supported local evidence

Local proof covers additive database migrations, tenant and ACL denial, resumable upload, durable PostgreSQL/Redis jobs, packaged Node 24 parser/query and check workers, worker restart recovery, large CSV and multi-sheet XLSX semantics, scoped learning, OpenAPI and client parity, MCP protocol behavior, and fail-closed feature and worker controls.

Not yet available

The following remain release blockers rather than supported behavior:
  • authenticated production browser acceptance;
  • named-host and deployed object-storage verification;
  • live external-provider and embedding execution;
  • registry publication of the candidate clients;
  • exact deployment identity and production-readiness proof; and
  • complete operations and policy-composition evidence across the release.
Until those proofs exist and the feature inventory changes status, do not depend on /quality or its candidate APIs in production.