# PCD measurements and limitations Measured 2026-09-16 on one NVIDIA L40S, FP16, PyTorch 2.14.0, Transformers 5.17.0, SDPA attention and reference PyTorch convolution. No optimized causal-conv1d, torch.compile, vLLM/SGLang or Jev service comparison. ## Release decision **Experimental inference-only release, not a production-quality or calibrated Jev substitute.** The latency goal is met on multi-field examples. The proposed accuracy non-inferiority goal is **not met**: use application-specific evaluation, an autoregressive fallback, or separately funded training before relying on it. No always-on service was deployed. ## Development and fresh audit The 12 development cases were used to debug/select prompts. Six additional cases were written after freezing the protocol and not used to retune it. Both sets are tiny, hand-authored, and not representative production data. Each method/case has one untimed warmup and three measured repetitions. Accuracy counts unique cases, not repeats. | Split | Method | Mean ms | Median ms | p95 ms | Field accuracy | Exact objects | Schema valid | |---|---|---:|---:|---:|---:|---:|---:| | development | token | 56.24 | 54.79 | 69.87 | 88.9% | 75.0% | 100.0% | | development | sequence | 60.47 | 55.50 | 85.64 | 77.8% | 66.7% | 100.0% | | development | ar_sequence | 557.41 | 538.50 | 659.95 | 97.2% | 91.7% | 100.0% | | development | ar_token | 565.64 | 536.43 | 691.47 | 91.7% | 83.3% | 83.3% | | audit | token | 57.34 | 55.08 | 70.98 | 72.2% | 33.3% | 100.0% | | audit | sequence | 58.09 | 56.13 | 69.60 | 66.7% | 33.3% | 100.0% | | audit | ar_sequence | 529.36 | 527.42 | 548.59 | 94.4% | 83.3% | 100.0% | | audit | ar_token | 559.70 | 546.89 | 629.99 | 94.4% | 83.3% | 100.0% | `ar_sequence` is greedy direct-answer JSON using original values and the same sequence-mode system/schema. `ar_token` uses the code catalog and an all-field JSON request; codes are strictly decoded, not repaired. Token PCD asks independent per-field questions, so its suffix prompt and distribution are not identical to joint AR. All use the same model, precision and backend. Empty thinking is explicit for these direct-answer modes. ## Synthetic scaling (one distinct input per row) | Case | Token PCD ms | Sequence PCD ms | AR JSON ms | AR/token ratio | Token / sequence / AR field accuracy | |---|---:|---:|---:|---:|---| | fields-4 | 53.31 | 52.00 | 748.99 | 14.1x | 100% / 100% / 100% | | fields-28 | 106.98 | 103.41 | 4875.30 | 45.6x | 100% / 100% / 100% | | fields-12 | 53.77 | 54.92 | 2099.88 | 39.1x | 100% / 100% / 100% | | enum-64 | 60.13 | 84.76 | 241.54 | 4.0x | 0% / 100% / 100% | | enum-255 | 144.99 | 346.77 | 270.86 | 1.9x | 0% / 100% / 100% | **High-cardinality warning:** token mode chose the wrong answer in both the 64- and 255-category probes. Sequence mode was correct on these inputs, but the 255-category sequence scorer was slower than AR. Do not interpret the high aggregate boolean accuracy as high-cardinality accuracy. The 28-field example is literal configuration extraction, not a broad semantic understanding benchmark. ## Incorrect constrained outputs (one per unique case) | Case | Method | Expected | Actual | |---|---|---|---| | routing-3 | token | `{"route": "north west", "service": "standard", "insured": true}` | `{"route": "north west", "service": "express plus", "insured": true}` | | routing-3 | sequence | `{"route": "north west", "service": "standard", "insured": true}` | `{"route": "north", "service": "express", "insured": false}` | | audit-3 | token | `{"route": "south", "service": "express plus", "insured": false}` | `{"route": "north west", "service": "express plus", "insured": false}` | | audit-3 | sequence | `{"route": "south", "service": "express plus", "insured": false}` | `{"route": "north", "service": "express", "insured": false}` | | audit-0 | token | `{"topic": "billing", "urgent": false, "refund": false}` | `{"topic": "technical", "urgent": false, "refund": false}` | | audit-2 | sequence | `{"route": "north", "service": "standard", "insured": true}` | `{"route": "north", "service": "express", "insured": false}` | | sentiment-2 | token | `{"sentiment": "neutral", "language": "Spanish", "question": true}` | `{"sentiment": "positive", "language": "English", "question": true}` | | audit-5 | token | `{"sentiment": "neutral", "language": "French", "question": false}` | `{"sentiment": "positive", "language": "English", "question": false}` | | sentiment-0 | sequence | `{"sentiment": "positive", "language": "English", "question": false}` | `{"sentiment": "positive", "language": "English", "question": true}` | | sentiment-3 | token | `{"sentiment": "neutral", "language": "English", "question": false}` | `{"sentiment": "positive", "language": "English", "question": false}` | | routing-0 | sequence | `{"route": "north west", "service": "express plus", "insured": true}` | `{"route": "north", "service": "express plus", "insured": false}` | | routing-1 | sequence | `{"route": "south", "service": "standard", "insured": false}` | `{"route": "north", "service": "express", "insured": false}` | | audit-1 | sequence | `{"topic": "technical", "urgent": true, "refund": true}` | `{"topic": "billing", "urgent": true, "refund": true}` | | audit-4 | sequence | `{"sentiment": "positive", "language": "Spanish", "question": false}` | `{"sentiment": "positive", "language": "Spanish", "question": true}` | | audit-4 | token | `{"sentiment": "positive", "language": "Spanish", "question": false}` | `{"sentiment": "positive", "language": "English", "question": false}` | | enum-64 | token | `{"category": "category-057"}` | `{"category": "category-017"}` | | enum-255 | token | `{"category": "category-248"}` | `{"category": "category-022"}` | ## Correctness and resource evidence - diagnostic: validation passed; maximum candidate-score differences against uncached reference {'token': 0.015625, 'sequence': 0.05997657775878906}; peak allocated VRAM 5.07 GiB; in-function time 94.24s. - stress: validation passed; maximum candidate-score differences against uncached reference {'token': 0.015625, 'sequence': 0.05997657775878906}; peak allocated VRAM 6.12 GiB; in-function time 79.74s. - Local tests exercise a tiny FP32 hybrid model; CUDA validation also tests the actual model in FP32 and FP16. Both single-token and multi-token convolution cache paths preserve the original prefix state. Score tolerances are 0.002 in FP32 and 0.10 in FP16; raw measured errors are retained, not assumed zero. - An early BF16 hidden-state check failed (maximum observed difference 0.25 with 0.15 absolute tolerance). We did not relax that BF16 gate; FP16 is the recommended tested runtime precision. Original stored BF16 weights remain unchanged. - Candidate probabilities are explicitly uncalibrated. No ECE, Brier-score calibration, learned confidence threshold, or proprietary Jev RLCD training is claimed. - Request timings include prompt tokenization, cached-schema lookup, prefill, cache forking, all microbatches, candidate scoring, and serialization; CUDA is synchronized. They exclude model loading, network, provisioning and output evaluation. Initial schema compilation and cold starts are not benchmarked separately. - p95 estimates are descriptive with very small samples. GPU clocks/allocator/runtime effects varied between exploratory jobs; the report uses the final fixed-source runs, not the fastest observed earlier latency. - Native reasoning was sampled during development on three cases with a 256-token cap: two truncated and one emitted fenced JSON. Those are not valid completed baseline results and are not used to claim PCD speedups or native-model incompetence. We avoided paying to repeat that capped experiment. - Frugal execution: serial single-L40S jobs, cached public weights, bounded runtimes, no H100, no training and no permanent deployment. In-function seconds are not a billing statement; startup/idle/CPU/storage charges are additional. ## Raw release evidence - [Diagnostic/audit JSON](../results/pcd/20260916-175747-benchmark-diagnostic-1f57e2.json) - [Scaling JSON](../results/pcd/20260916-175954-benchmark-stress-d12f1b.json) Raw files include every measured output, score, token/branch count, version and inference-source hash. The selected prompt was not changed after the fresh audit. Earlier iterations are development artifacts and are not release performance claims.