Run the unchanged benchmark as a separate advisory job while #78 is open. Keep functional coverage and other checks required, and retain the 2 GB heap and 300-second protocol limit. Record #78's obligation to restore the gate. Refs: #75, #78
21 KiB
Mobile leave-planner benchmark
Composable optimization rules: current verification status
The #70–#75 implementation uses explicit active optimization configurations in the benchmark fixtures. The latest full functional coverage run passes 652 tests across 34 files under the 2 GB Node heap limit, excluding the separate benchmark. This does not establish a passing performance protocol.
MR pipeline 320 passed all 652 functional tests in 109.43 seconds, then the seven-scenario protocol timed out at its unchanged 300-second limit without an out-of-memory failure. In an earlier diagnostic run, the loaded-year warm-up took 184.36 seconds and its first measured repetition began about 251.78 seconds into the protocol. Those diagnostic timings precede the final period-projection cache and the planned original-decision-trace removal, and are not a result for either.
At the user's request, #75 finishes on saved-rule migration and functional
verification. #78 owns benchmark verification and implementation of further
optimizations when needed, after #76 removes tracing and #77 verifies the retained
behavior. All seven scenarios, one warm-up and two measured repetitions per
scenario, the 2 GB heap limit and the 300-second timeout remain required for #78.
At the user's request, CI runs the benchmark in a separate job with
allow_failure: true while #78 is open. Functional coverage tests and all other
required checks remain blocking. #78 must restore the required benchmark gate
after the unchanged protocol passes. A green pipeline during this exception
does not establish a passing benchmark. No automatic elapsed-time planning
cutoff is introduced.
The results below document earlier implementations and protocols. They are historical measurements rather than the performance result for composable rules.
Reference-phone protocol
Open the deployed application at /benchmark on the reference phone. Tap
Run benchmark, wait for the seven scenarios to finish, then tap Copy
report and save the tab-separated output with the date and build identifier.
The harness measures preparation, planning-to-outcome latency, later trace
delivery, and total wall time separately. It also reports planning work, trace
capture, trace materialization, and trace transfer independently, plus the
largest extra delay of a 50 ms UI heartbeat. It records the outcome, last
completed quality stage, placement count, determinism, trace status, and trace
candidate count. PlanningIncomplete is not a conflict: its last completed
stage and incumbent are retained.
Protocol
- Record phone model, OS version, browser/version, application build, and battery level.
- Use a cool phone, unplugged, with other applications closed. Record whether the device became warm or throttled during the run.
- Tap Run benchmark once. It first runs every scenario as an unreported warm-up, then records three measured repetitions per scenario. Do not interact with the page while a run is active. The first measured result is the determinism baseline; later repetitions report whether they match it.
- The scenarios include two and three calendar years with separately dated annual CP and RTT balances. Let these runs finish or cancel explicitly; there is no automatic calculation cutoff and no validated duration ceiling for multi-year horizons. The historical annual targets and release exception below do not establish multi-year performance.
- Keep the copied report unchanged, including any incomplete outcomes or hardware limitations. No mobile-device measurement is claimed by the source tree until a human records it here.
The browser page uses the same createPlanningCalculation(request) context as
the calendar and adds no solver or backend dependency. It omits deadlineMs.
Trace delivery is reported separately and cannot replace a valid planning outcome.
planning_ms is the observed latency until the outcome is usable;
planning_work_ms excludes trace capture. trace_capture_ms accounts for
instrumentation during planning, trace_materialization_ms accounts for
building the stable explanation and finishing the trace after the outcome.
trace_transfer_ms measures worker transfer and receiver assembly work;
trace_delivery_ms measures elapsed time until the complete trace is ready. max_ui_delay_ms includes
planning, trace delivery, and report preparation. Trace readiness has no
elapsed-time ceiling and never permits truncating actual candidate evaluations.
peak_main_thread_heap_bytes is Chromium's available performance.memory
sample, taken alongside the 50 ms UI heartbeat. A blank value means this API is
unavailable. This measures the page's JS heap, excludes the worker heap, and is
not a claim about total application memory. For a browser profile, record the
isolated Chrome process-tree RSS separately during planning and subsequent
trace delivery, including the worker and renderer. State the sampling interval,
CPU throttling, viewport, and whether a warm-up preceded each measurement.
For cancellation measurements, use the normal calendar with the same saved two- or three-year dates and balances. Cancel during preparation and an active optimization phase, using the keyboard in one run. Record request-to-outcome latency separately from request-to-artifacts latency; verify the partial-plan status or absence of a plan, one planning request, no placements before the outcome, and continued acknowledged delivery of the original trace. Confirm that editing an input remains possible during active worker CPU work. These measurements describe that device and workload; they impose no latency limit.
Hardware status
No reference-phone run has been recorded. The project owner elected to close the release tickets on 2026-09-25 using the emulated run below, accepting its six-to-seven-second loaded planning time and the observed post-result trace UI pause. This is an explicit release exception, not a reference-phone measurement or proof of physical-device responsiveness.
Long-horizon calendar verification — 2026-10-05
Working implementation of issue #59 based on
2edf2004f56f8a453a691846f3c0e6d960e27911, using the production browser build,
native Chromium 149.0.7827.55 Web Workers, and an AMD Ryzen 7 7800X3D desktop.
The host ran Debian 13 under WSL2 (kernel 6.18.33.2-microsoft-standard-WSL2).
Headless browser emulation used 412×915 pixels, scale 2.625, touch enabled,
reduced motion, and 4× CPU throttling. Each scenario used a fresh browser and
one cold run, with no warm-up. Other validation commands were kept out of these
runs. This is desktop evidence, not a reference-phone performance claim.
The saved dates covered 2024-01-01 through 2025-12-31 or 2026-12-31 inclusive.
Each year supplied 20 CP and 10 RTT days available January 1 and expiring
December 31, matching the new benchmark scenarios. No technical deadline was
supplied. Both completed plans finished through the dates objective, consumed
all balances, and displayed no intermediate placements. The calendar appeared
before the full trace; its placement dates were identical after trace readiness,
and each flow sent exactly one planning request. All 1,949 and 3,956 trace
packets respectively were acknowledged, carrying 122,621 and 249,234 candidate
records without truncation. Heartbeats continued during the producing runs.
| Horizon | Calendar usable | Later trace delivery | Placements | Maximum UI timer delay before calendar / during trace | Peak page heap before calendar / during trace | Peak Chrome process-tree RSS before calendar / during trace |
|---|---|---|---|---|---|---|
| Two years | 5.08 s | 2.55 s | 60 | 1.47 s / 0.70 s | 22.7 / 314.6 MiB | 1,143.2 / 1,606.8 MiB |
| Three years | 10.86 s | 4.14 s | 90 | 2.11 s / 1.48 s | 22.9 / 597.6 MiB | 1,386.8 / 2,217.8 MiB |
A 50 ms page timer measured extra event-loop delay and sampled available main-thread JS heap. RSS was sampled every 500 ms across that isolated Chrome process tree, including browser, renderer, worker, and auxiliary processes; it is resident process memory, not worker JS heap. The timer delays include initial multi-year calendar rendering and later artifact work. These pauses and memory costs remain visible limitations; no duration or responsiveness threshold is inferred from these samples.
Two additional cold three-year runs used the same balances in dark mode.
Cancellation during preparation returned no plan and no obligation conflict.
Keyboard cancellation during the holiday-week pass retained a valid 90-placement
incumbent marked Calcul annulé — plan partiel, with at-risk as the last
completed stage. An editable input accepted focus while the worker remained
active. Both flows kept a 44 px Cancel control, fit the viewport without
horizontal overflow, and delivered the original trace after their outcome.
| Cancellation point | Request to outcome | Request to complete artifacts | Retained placements |
|---|---|---|---|
| Calendar preparation | 21.9 ms | 874.8 ms | 0 |
| Active holiday-week optimization, keyboard | 21.5 ms | 1,431.7 ms | 90 |
No browser errors occurred. Cancellation issued one cancel request and no additional calculation. Separate tests cover late artifact errors and stale outcomes, explanations, and traces, preserving the latest displayed plan. Elapsed-time tests advance clocks beyond ten seconds without waiting in CI, including inside the real bundled worker. The seven-scenario benchmark also checks complete traces and repeated outcome determinism for both longer horizons.
Indicative wait and complete horizon flow — 2026-10-05
Working implementation of issue #60, preserving the existing #59 worktree
changes, based on 2edf2004f56f8a453a691846f3c0e6d960e27911. Production build,
native Chromium 149.0.7827.55 workers, AMD Ryzen 7 7800X3D, Debian 13.6 under
WSL2 (kernel 6.18.33.2-microsoft-standard-WSL2). Headless emulation used 412×915
pixels, scale 2.625, touch, reduced motion and 4× CPU throttling. External font
downloads were blocked. No other build or test workload ran during measurement.
These are desktop samples, not physical-phone results or accuracy guarantees.
Each scenario used a fresh browser. The application's initial calculation covered one saved day with no balances or rules; that small completed plan was the session's sole prior calibration sample. No target-scenario warm-up was performed. Through the actual settings and balance controls, the run then selected January 1, 2024 through the relevant December 31 inclusive and supplied 20 CP plus 10 RTT days per year, with annual availability and expiry dates. The saved dates and balances were checked. All edits retained the initial calendar and issued no calculation; one explicit Calculate click submitted the complete target configuration.
Calendar wait was measured from the captured Calculate click to DOM observation
of the displayed final placements with the loading overlay removed. This
includes request scheduling, preparation and rendering, unlike a core-work
measurement. The later trace interval starts at that calendar endpoint and ends
at the worker's terminal artifact delivery. The application's own calibration
uses its request timestamp and afterNextRender endpoint as described in
planning-duration.md.
| Horizon | Range before Calculate (with one local sample) | Observed calendar wait | Additional trace delivery | Range after completion | Placements |
|---|---|---|---|---|---|
| One year | 1–4 s | 1.74 s | 1.23 s | 1–5 s | 30 |
| Two years | 2–16 s | 5.08 s | 2.88 s | 2–16 s | 60 |
| Three years | 7–44 s | 11.03 s | 5.60 s | 6–39 s | 90 |
All three outcomes completed through dates, with no intermediate placements,
implicit deadline, extra request or date movement. The three-year result
finished after ten seconds. Calendars and estimates remained identical during
the later trace interval. All 674, 1,948 and 3,959 trace packets respectively
were acknowledged. These isolated samples establish behavior, not a validated
numerical prediction error or duration ceiling. The rough model before any
local sample is separately documented; it does not assume linear growth with
the number of years.
Two dark-mode native-worker cancellation flows covered the same manually submitted three-year configuration. Holding the new worker's script download until the explicit Cancel click made cancellation before any valid plan reproducible: no plan or conflict was presented, and the indicative 7–46 s range was unchanged. Its empty original trace was already available when the calendar rendered. A keyboard cancellation in active holiday-week optimization retained 90 placements, displayed Calculation cancelled — partial plan, and left its 7–44 s range unchanged through later trace delivery. Both Cancel controls were 44 px high, and balance inputs accepted focus while the calculation was active. The controlled startup download is not cancellation-latency evidence.
A separate three-year run completed in 10.96 s while native trace-update events were held back. The calendar was usable with an explicitly loading trace; injecting a subsequent worker transport error preserved all 90 placements and the newly refined 6–38 s estimate. Controlled-clock tests additionally keep a trace unresolved for five minutes after a twenty-second calendar wait, verifying that subsequent delivery/error time contributes no calibration sample.
No browser errors or horizontal overflow occurred. The annual flow also fit
375, 768 and 1440 px widths. The estimate uses the existing secondary-text
theme token and is associated with Calculate through aria-describedby.
Settings restoration, stale calculations, rule/balance edits, long-worker
liveness, and cancellation history are also exercised by the integration suite.
Emulated browser run — 2026-09-25
Chromium 149.0.7827.55 on an AMD Ryzen 7 7800X3D (x86_64), headless at 412×915 pixels, device scale 2.625, touch enabled, and 4× CPU throttling. One full warm-up was excluded, followed by three measured repetitions. All five scenarios returned within ten seconds of planning time and complete results matched across repetitions. The loaded year took 5.74–6.17 seconds, above the five-second target. A 100 ms main-thread timer observed a maximum extra delay of 5,950.9 ms. Worker event timestamps place the four loaded-year pauses of 4.0–6.0 seconds after their planning results, around delivery of full trace artifacts. This desktop emulation does not establish reference-phone performance or responsiveness.
scenario repetition preparation_ms planning_ms trace_delivery_ms wall_ms outcome last_stage placements deterministic trace_status trace_candidates
representative year 1 0.0 443.7 624.7 1069.1 Plan dates 30 baseline ready 29451
representative year 2 0.1 490.0 609.3 1099.4 Plan dates 30 yes ready 29451
representative year 3 0.1 465.5 614.2 1079.8 Plan dates 30 yes ready 29451
typed obligations 1 0.0 157.0 261.2 418.2 Plan dates 10 baseline ready 13103
typed obligations 2 0.0 162.6 255.6 418.2 Plan dates 10 yes ready 13103
typed obligations 3 0.0 163.4 258.2 421.6 Plan dates 10 yes ready 13103
chained earned entitlements 1 0.0 45.3 1.0 47.0 Plan dates 3 baseline ready 27
chained earned entitlements 2 0.1 44.5 0.9 45.5 Plan dates 3 yes ready 27
chained earned entitlements 3 0.1 44.4 0.8 45.4 Plan dates 3 yes ready 27
deliberate conflict 1 0.0 92.8 0.9 93.8 PlanningConflict 0 baseline ready 366
deliberate conflict 2 0.0 44.7 0.2 44.9 PlanningConflict 0 yes ready 366
deliberate conflict 3 0.0 43.5 0.6 44.1 PlanningConflict 0 yes ready 366
loaded year 1 0.3 5901.5 5677.7 11581.0 Plan dates 144 baseline ready 865117
loaded year 2 0.8 6167.5 7582.4 13752.4 Plan dates 144 yes ready 865117
loaded year 3 0.6 5738.3 6219.5 11962.5 Plan dates 144 yes ready 865117
Exact-trace verification — 2026-10-03
Working changes based on 0ce33cc9d6fe7381965a9162571abe205d0b7350,
Chromium 149.0.7827.55 on the same AMD Ryzen 7 7800X3D desktop,
412×915 pixels, device scale 2.625, touch enabled, and 4× CPU throttling.
One warm-up preceded three recorded repetitions. This is desktop emulation;
no physical-phone result is claimed, and the earlier release exception is not
used as evidence for this implementation.
All five scenarios matched their pre-change outcomes in a separate core comparison, including traced versus untraced runs. Every repeated browser outcome matched its first repetition. Each loaded run returned 144 placements and retained all 868,919 actual trace candidate records with a ready trace.
Loaded planning work took 5.81–6.18 seconds, below its ten-second budget. Observed outcome latency was 10.36–10.85 seconds, including trace capture; this latency is distinct from the planning-work budget. Complete trace delivery followed 24.09–25.10 seconds later. No evaluation was truncated to meet a timing target. The worst extra delay of the 50 ms heartbeat was 946.8 ms across the loaded runs, including report preparation, compared with 4.65–5.49 seconds before acknowledged delivery in the same final-trace harness.
The worker stages snapshots during planning and delivers at most sixteen 64-operation packets before waiting for receiver acknowledgements. The receiver applies bounded batches, releases consumed packets, and yields between tasks. The worker stops after its terminal packet; benchmark helpers keep completed trace references out of the next repetition. These changes avoid the observed message-flood pauses and repeated-run renderer crashes. The complete trace still requires substantial memory: a separate unthrottled core profile retained approximately 1.06 GB of heap for the loaded trace. These results do not establish physical-device performance.
A separate rendered-calendar run at the same emulated mobile size used twelve balances of twelve days each. It produced 144 placements, 113 gained non-working days, and a longest work sequence of five. Full trace navigation and date comparison completed without browser errors. Maximum extra heartbeat delays were 487.4 ms during planning and calendar rendering, 380.8 ms during trace exploration, and 540.3 ms during comparison. Keyboard navigation, dialog focus restoration, background comparison without an implicit modal, and responsive widths of 375, 768, and 1440 pixels also passed.
The final suite passes 201 tests across 23 files, application and specification type checks, ESLint, and the production build. Build warnings remain for the 2.55 MB initial bundle (2.50 MB warning budget) and 13.41 kB calendar stylesheet (8.00 kB warning budget); hard budgets are not exceeded.
scenario repetition preparation_ms planning_ms trace_delivery_ms wall_ms outcome last_stage placements deterministic trace_status trace_candidates max_ui_delay_ms planning_work_ms trace_capture_ms trace_materialization_ms trace_transfer_ms
representative year 1 0.0 622.4 716.9 1340.3 Plan dates 30 baseline ready 29296 18.2 265.0 256.4 7.4 166.4
representative year 2 0.0 593.5 695.8 1289.4 Plan dates 30 yes ready 29296 14.5 255.3 233.8 7.5 184.1
representative year 3 0.0 592.0 721.0 1313.2 Plan dates 30 yes ready 29296 11.5 256.1 234.4 9.2 191.8
typed obligations 1 0.1 259.7 324.3 584.6 Plan dates 10 baseline ready 13005 11.8 106.8 56.6 3.7 70.2
typed obligations 2 0.1 266.2 330.6 596.9 Plan dates 10 yes ready 13005 11.7 110.6 60.3 4.1 81.9
typed obligations 3 0.1 259.6 341.8 601.5 Plan dates 10 yes ready 13005 15.7 112.3 54.7 4.0 75.9
chained earned entitlements 1 0.1 94.5 3.0 97.6 Plan dates 3 baseline ready 15 0.8 3.4 1.3 0.9 0.3
chained earned entitlements 2 0.0 100.2 2.5 102.7 Plan dates 3 yes ready 15 0.2 3.6 1.4 0.7 0.1
chained earned entitlements 3 0.0 96.0 1.5 97.5 Plan dates 3 yes ready 15 0.1 3.2 1.5 0.7 0.1
deliberate conflict 1 0.0 94.4 2.4 96.8 PlanningConflict 0 baseline ready 4 0.2 0.9 0.8 0.3 0.2
deliberate conflict 2 0.0 93.2 1.0 95.0 PlanningConflict 0 yes ready 4 0.2 1.1 0.5 0.2 0.0
deliberate conflict 3 0.1 92.8 1.6 95.2 PlanningConflict 0 yes ready 4 1.4 1.0 0.7 0.2 0.1
loaded year 1 0.0 10852.1 25102.6 36002.9 Plan dates 144 baseline ready 868919 946.8 6014.8 4730.6 948.3 8862.1
loaded year 2 0.1 10364.4 24816.9 35216.6 Plan dates 144 yes ready 868919 791.2 5813.1 4440.6 296.7 8693.3
loaded year 3 0.1 10807.8 24089.4 34933.6 Plan dates 144 yes ready 868919 721.7 6176.0 4526.6 978.7 8408.0