ci: make the deferred planning benchmark non-blocking
Run the unchanged benchmark as a separate advisory job while #78 is open. Keep functional coverage and other checks required, and retain the 2 GB heap and 300-second protocol limit. Record #78's obligation to restore the gate. Refs: #75, #78
This commit is contained in:
1 parent
6811142299
commit
98e87fecee
5 files changed
+31
-14
No files matched your search
+9
-2
@@ -85,8 +85,6 @@ test:
|
||||
- bun scripts/generate-langs.js
|
||||
# ponytail: one worker bounds peak RAM; increase concurrency after adding runner memory.
|
||||
- bunx vitest run --coverage --maxWorkers=1 --exclude='src/app/benchmark/**'
|
||||
# Keep the complete benchmark outside coverage instrumentation.
|
||||
- bunx vitest run src/app/benchmark/leave-planner-benchmark.spec.ts --maxWorkers=1
|
||||
artifacts:
|
||||
when: always
|
||||
paths:
|
||||
@@ -97,6 +95,15 @@ test:
|
||||
path: coverage/cobertura-coverage.xml
|
||||
expire_in: 7 days
|
||||
|
||||
# shortcut: benchmark failures are advisory while #78 is open; restore the gate when #78 passes.
|
||||
benchmark:
|
||||
extends: test
|
||||
allow_failure: true
|
||||
script:
|
||||
- bun scripts/generate-langs.js
|
||||
- bunx vitest run src/app/benchmark/leave-planner-benchmark.spec.ts --maxWorkers=1
|
||||
artifacts: null
|
||||
|
||||
# -- Build Stage --
|
||||
|
||||
build:
|
||||
|
||||
+2
-2
@@ -4,7 +4,7 @@
|
||||
|
||||
Blocked by: #74
|
||||
|
||||
Les exigences de performance sont transférées à #78. La clôture de ce ticket dépend de la migration et de la validation fonctionnelle, pas du passage du protocole de performance. Les limites du benchmark et les contrôles CI restent inchangés.
|
||||
Les exigences de performance sont transférées à #78. La clôture de ce ticket dépend de la migration et de la validation fonctionnelle, pas du passage du protocole de performance. Les limites du benchmark restent inchangées. À la demande de l’utilisateur, son job CI est séparé et non bloquant tant que #78 est ouvert ; #78 doit rétablir ce contrôle obligatoire après réussite du protocole. Les tests fonctionnels et les autres contrôles CI restent obligatoires.
|
||||
|
||||
## What to build
|
||||
|
||||
@@ -25,7 +25,7 @@ Livraison intégrée de l’ADR 0008 sur la branche commune aux tickets #70, #71
|
||||
- [x] Les explications et analyses contrefactuelles utilisent le même ordre actif que le calcul. Modifier les règles durant le calcul ou pendant la livraison des explications ne change pas ses résultats. Les anciens objectifs implicites n’interviennent plus dans aucune comparaison ou protection.
|
||||
- [x] Les fixtures et scénarios de benchmark qui dépendaient d’une optimisation implicite demandent explicitement les préréglages nécessaires. Les tests d’absence de règles n’utilisent aucun défaut implicite pour retrouver l’ancien résultat.
|
||||
- [x] Exécuter les suites fonctionnelles concernées de règles, interface, moteur, worker réel, explication, annulation et estimation de durée, ainsi que les contrôles habituels de compilation et de qualité. Préserver les assertions de résultat et de déterminisme des fixtures, y compris les horizons longs. L’exécution complète du protocole de benchmark, la réactivité mesurée, les durées, la consommation mémoire et les optimisations nécessaires relèvent de #78 et ne conditionnent plus la clôture de #75. Ne pas introduire de délai d’arrêt automatique.
|
||||
- [x] La vérification fonctionnelle commune des tickets #70, #71, #72, #73, #74, #75 est la condition de clôture fonctionnelle de ce ticket. Documenter les choix finaux du contrat composable ; ne laisser ni double moteur, ni chemin historique implicite, ni adaptateur provisoire ajouté pour rendre les étapes indépendamment livrables. Les limites de performance mesurées et leur amélioration sont suivies dans #78 ; les exigences habituelles de CI et de mise en service ne sont pas abaissées.
|
||||
- [x] La vérification fonctionnelle commune des tickets #70, #71, #72, #73, #74, #75 est la condition de clôture fonctionnelle de ce ticket. Documenter les choix finaux du contrat composable ; ne laisser ni double moteur, ni chemin historique implicite, ni adaptateur provisoire ajouté pour rendre les étapes indépendamment livrables. Les limites de performance mesurées et leur amélioration sont suivies dans #78 ; les tests fonctionnels et les autres contrôles CI restent obligatoires, avec la seule exception temporaire du job de benchmark décrite ci-dessus.
|
||||
|
||||
**Vérification de bout en bout :** Partir d’une ancienne sauvegarde contenant des règles personnalisées et désactivées, charger la nouvelle version, vérifier le premier calcul, modifier ou supprimer des préréglages, recharger et constater la conservation exacte des choix.
|
||||
|
||||
|
||||
+2
-1
@@ -18,6 +18,7 @@ If the measured baseline already passes, keep the engine unchanged and record th
|
||||
|
||||
- [ ] Keep the existing seven benchmark scenarios and inputs, including the loaded year, typed joint obligations, chained earned entitlements and two- and three-year horizons. Keep explicit optimization configurations, quotas, expected outcomes, placement counts, completed-stage and determinism assertions.
|
||||
- [ ] Run the real existing benchmark protocol: one warm-up and two measured repetitions per scenario, outside coverage instrumentation, with a 2 GB Node heap limit and the 300-second timeout for the full protocol. All scenarios and repetitions must finish and pass; a timeout, skipped scenario or partial report does not satisfy this ticket.
|
||||
- [ ] Restore the benchmark as a required CI check after the unchanged protocol passes. While this ticket is open, the user approved a separate automatic benchmark job with `allow_failure: true` so its failure does not block functional delivery. Passing CI during this exception does not satisfy benchmark acceptance or justify closing this ticket.
|
||||
- [ ] After #76 removes tracing and #77 verifies the retained behavior, benchmark outcome and explanation delivery without obsolete trace counters, trace transfer or trace-acknowledgement waits. Do not reintroduce tracing as profiling infrastructure.
|
||||
- [ ] When the baseline fails, profile CPU, allocation and repeated planning work; apply optimizations and rerun the unchanged protocol until it passes. Preserve planning semantics, rule ordering, joint obligations, earned-entitlement chronology, final-plan quality and deterministic tie-breaking. Use behavior-level or differential tests to demonstrate that caching, deduplication or pruning is valid across the relevant configurations.
|
||||
- [ ] Preserve responsive browser interaction, explicit cancellation, heartbeat monitoring and best-valid-partial-plan recovery during representative and long-horizon calculations. Keep the existing duration-estimation contract and recalibrate estimates only when measured results require it.
|
||||
@@ -30,7 +31,7 @@ If the measured baseline already passes, keep the engine unchanged and record th
|
||||
|
||||
The previous real protocol timed out at 300 seconds without an out-of-memory failure. The loaded-year warm-up took about 184 seconds and its first measured repetition started about 252 seconds into the protocol. Original trace recording and repeated allocation previews were observed costs; the benefit of trace removal has not yet been measured. These are starting points for profiling, not permission to change optimization quality.
|
||||
|
||||
This ticket authorizes verification and implementation of performance improvements when needed. Functional migration remains in #75, trace removal in #76 and trace-removal functional verification in #77. Existing CI benchmark limits remain unchanged.
|
||||
This ticket authorizes verification and implementation of performance improvements when needed. Functional migration remains in #75, trace removal in #76 and trace-removal functional verification in #77. Existing benchmark limits remain unchanged; only the benchmark job's blocking status is temporarily waived at the user's request until this ticket restores it.
|
||||
|
||||
## Blocked by
|
||||
|
||||
|
||||
@@ -7,18 +7,24 @@ the benchmark fixtures. The latest full functional coverage run passes 652 tests
|
||||
across 34 files under the 2 GB Node heap limit, excluding the separate benchmark.
|
||||
This does not establish a passing performance protocol.
|
||||
|
||||
The most recent seven-scenario protocol attempt timed out at its unchanged
|
||||
300-second limit without an out-of-memory failure. The loaded-year warm-up took
|
||||
184.36 seconds; its first measured repetition began about 251.78 seconds into
|
||||
the protocol. These measurements precede the final period-projection cache and
|
||||
the planned original-decision-trace removal, and are not a result for either.
|
||||
MR pipeline 320 passed all 652 functional tests in 109.43 seconds, then the
|
||||
seven-scenario protocol timed out at its unchanged 300-second limit without an
|
||||
out-of-memory failure. In an earlier diagnostic run, the loaded-year warm-up took
|
||||
184.36 seconds and its first measured repetition began about 251.78 seconds into
|
||||
the protocol. Those diagnostic timings precede the final period-projection cache
|
||||
and the planned original-decision-trace removal, and are not a result for either.
|
||||
|
||||
At the user's request, #75 finishes on saved-rule migration and functional
|
||||
verification. #78 owns benchmark verification and implementation of further
|
||||
optimizations when needed, after #76 removes tracing and #77 verifies the retained
|
||||
behavior. All seven scenarios, one warm-up and two measured repetitions per
|
||||
scenario, the 2 GB heap limit and the 300-second timeout remain required. No
|
||||
automatic elapsed-time planning cutoff or weaker CI gate is introduced.
|
||||
scenario, the 2 GB heap limit and the 300-second timeout remain required for #78.
|
||||
At the user's request, CI runs the benchmark in a separate job with
|
||||
`allow_failure: true` while #78 is open. Functional coverage tests and all other
|
||||
required checks remain blocking. #78 must restore the required benchmark gate
|
||||
after the unchanged protocol passes. A green pipeline during this exception
|
||||
does not establish a passing benchmark. No automatic elapsed-time planning
|
||||
cutoff is introduced.
|
||||
|
||||
The results below document earlier implementations and protocols. They are
|
||||
historical measurements rather than the performance result for composable rules.
|
||||
|
||||
@@ -279,5 +279,8 @@ Ticket #75 covers saved-rule migration and functional integration of #70–#75.
|
||||
Performance acceptance and any further optimization are tracked in #78, after
|
||||
original-decision-trace removal in #76 and functional verification in #77. The
|
||||
seven-scenario benchmark, its warm-up and two measured repetitions, the 2 GB heap
|
||||
limit and the 300-second timeout remain unchanged. Completing the functional
|
||||
migration does not establish that these performance gates pass.
|
||||
limit and the 300-second timeout remain unchanged. CI runs the benchmark as a
|
||||
separate non-blocking job at the user's request while #78 is open; functional
|
||||
tests and other checks remain required. #78 must restore the required benchmark
|
||||
gate after the protocol passes. Completing the functional migration or passing
|
||||
CI during this exception does not establish a passing benchmark.
|
||||
Reference in new issue
Block a user