Compare commits

...

12 Commits

Author SHA1 Message Date
3429bce8d0 let domains author their own pass-2 estimator as formulas, not Python
Pass 2 previously only knew one estimator: a hardcoded physics model that
matches dimensions literally named platform/actuator/energy_storage. Any
domain outside that shape (e.g. archery) got all-zero estimates and failed
every combo. Domains can now declare free variables and per-metric formulas
as data instead; a safe AST-based evaluator (engine/formula.py, no eval())
resolves declared entity properties via dep(key, constraint_type) and
generalizes the existing hand-nested mass-budget search into an N-variable
recursive optimizer. Fully additive -- the legacy platform/actuator/
energy_storage path is untouched and still runs unchanged for every domain
that declares no formulas.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-16 21:28:15 -05:00
3795a7e826 rename estimator, skip unused cargo calc, fix six review bugs, close three seed guardrail holes
Rename: _stub_estimate -> _estimate_physics (it's a real deterministic
physics engine now, not a stub) and estimation_method "stub" ->
"physics_calc" to match the value pass 3 already used, for consistency
between the raw-estimate and scored-metric tables. Also skip the
cargo_capacity/cargo_capacity_kg arithmetic entirely in
_raw_physics_from_masses for domains that score neither and don't need
it as cost_efficiency's $/(kg·m) denominator either -- real but modest
savings on the ~11,000-eval-per-combo optimizer hot path (a separate
log1p-caching attempt was tried and reverted: it measured SLOWER, not
faster -- the extra dict lookup cost more than the two math.log1p calls
it avoided).

Six bugs found by a full-codebase review agent, verified individually:

- pipeline.py: LLM rate-limit retry called review_plausibility() with
  domain.metric_bounds instead of domain, crashing the whole pipeline
  run on any retry (every provider immediately accesses domain.name/
  .metric_bounds on that arg).
- _explore_result.html: mass-bar width divided by total_mass with no
  zero guard; biological/ambient actuators can legitimately have 0 mass
  floors, so an all-zero slider combination 500'd the explore endpoint.
- routes/pipeline.py: if init_db/Repository(conn) raised before
  repo/conn were assigned, the except/finally handlers referencing them
  raised UnboundLocalError, silently swallowed by bare except/pass --
  a bad PHYSCOM_DB path left a run stuck at status=pending forever with
  no diagnostic. conn/repo now init to None and are guarded before use;
  the truly-unreachable-DB case at least logs server-side now.
- repository.py: update_combination_status's downgrade guard protected
  scored/llm_reviewed/*_fail but not a write of "valid" -- pass 1
  re-running for a different domain against an already-reviewed combo
  silently reverted its status back to "valid", erasing the review
  signal. Verified directly: marked a combo reviewed, re-ran pass 1,
  status held.
- pipeline.py: cost_efficiency's operating-cost term fell back to
  ground rolling-resistance physics (effective_k_med or ...["ground"])
  for media with no resistance model (space), instead of skipping the
  term the way range_fuel explicitly does two lines above. Every scored
  interplanetary_travel combo got a cost_efficiency computed from
  ground physics applied to a spacecraft. Now reports amortized/upfront
  cost only for such media -- an honest partial answer.
- pipeline.py: `if min_accel and specific_thrust:` used truthiness
  instead of `is not None` -- dep_value() legitimately returns 0.0 for
  a declared floor of zero (Spaceship declares min_effective_accel=0),
  masking a real requirement as "undeclared."

Three seed-data guardrail holes, matching LOGIC DOCS/002's "missing
floor is a silent hole" pattern:

- constraint_resolver.py: CATEGORY_SEVERITY had no entry for the
  "material" category, so Nuclear Thermal Drive/Nuclear Fuel's
  radiation_shielding requirement defaulted to a non-blocking "warn"
  nothing in the catalog ever satisfies. Added material -> block.
  Consequence, verified: every nuclear combo across all domains now
  correctly fails pass 1, since nothing currently provides shielding --
  the accurate state given the catalog gap, not a regression.
- transport_example.py: Submarine had a mass range_min but no
  range_max, unlike its sibling water platform -- _decide_masses skips
  its entire structural-feasibility search when p_max is None. Added a
  20,000,000kg ceiling (small submersible to large ballistic-missile
  class).
- transport_example.py: Amphibious Vehicle declared no medium requires
  at all, so it vacuously satisfied every domain's medium constraint
  including space-only interplanetary_travel. Added medium=ground
  (the current requires model has no OR semantics for "ground or
  water," so this is a real tradeoff -- it can no longer participate in
  maritime_shipping either, losing the water half of "amphibious").
  Verified: interplanetary_travel's pass-2-estimated count dropped from
  33 to 3, and all 3 remaining are genuinely Spaceship-based; the ~30
  removed were confirmed to be Amphibious Vehicle's vacuous passes.

Logged the GPU-batching-for-the-optimizer discussion (why it doesn't
fit at current scale, what threshold would change that, what it would
actually require) as LOGIC DOCS/003 for future reference.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-16 17:33:22 -05:00
81b36e6bbe extend the drag term to air medium -- phi4 caught what the ground-only fix missed
Running a real phi4 pass-4 review pass surfaced a Rotorcraft + Gas
Turbine combo with a "speed" of 3,127 m/s (Mach 9), and the model's own
review text called it out directly: "unrealistic for urban commuting,
likely indicating an error." The reasoning for leaving air out of
DRAG_POWER_COEFF_BY_MEDIUM ("its L/D-based cruise model is already a
reasonable velocity-roughly-linear approximation") was wrong for the
same reason ground was wrong before: L/D-based drag force is also
roughly velocity-independent within a design cruise band, so it's still
just a mass-proportional constant with no v^2 term pushing back, and
inverting power/resistance for achieved speed had no ceiling there
either.

Added an aircraft-like reference cross-section (0.5 * rho_air * Cd(~0.2)
* frontal_area(~3.5 m^2)) to DRAG_POWER_COEFF_BY_MEDIUM for "air",
reusing the same closed-form cubic solve already built for ground.
Confirmed: the same combo's speed dropped from 3,127 m/s to 125.9 m/s
(a fast but physically plausible rotorcraft cruise), and a second phi4
pass-4 run on the corrected data no longer flags it -- the review now
discusses the speed score as a genuine strength instead of an apparent
error.

Water hull drag remains a known, unaddressed gap (would need its own
reference, not a car's or aircraft's frontal area).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 21:37:28 -05:00
fb38093e6c fix cargo_capacity direction: storage competes with cargo, doesn't pad it
cargo_capacity was based on floor_total (platform+actuator+storage), so
more battery/fuel made cargo capacity go UP -- backwards. Real deadweight
tonnage is a fixed allowance sized off the vessel's own empty (lightship)
mass -- hull + machinery, not fuel -- and fuel and cargo then SHARE that
one allowance: more fuel bunkered means less room left for cargo. Cargo
capacity is now (platform+actuator)*ratio - storage_mass, floored at 0,
so reducing battery/fuel now correctly frees up cargo room instead of
shrinking it.

Confirmed on combo #876: at the battery's declared 9kg floor, cargo is
6.3kg; increasing to 15kg drops it to 0.3kg, to 30kg+ drops it to 0 --
storage size and cargo capacity now trade off in the right direction.

Re-ran all five domains: urban_commuting's insufficient_structure count
went from 1 to 6 passing combos, since cargo_capacity's changed
incentives shifted where the optimizer lands relative to the structural
cap -- checked each one, all are 0.007-0.033% over the cap, the same
grid-search rounding-noise magnitude already established as acceptable,
just now triggered on a few more combos.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 20:35:48 -05:00
f786f3da79 derive cargo_capacity from the actual optimized build, not the declared-floor sum
cargo_capacity/cargo_capacity_kg were computed from the SUM of each
entity's own declared minimum mass (the same floor pass 1 checks for
legality) -- completely disconnected from the platform/actuator/storage
masses _decide_masses actually optimizes and every other metric
(power_density, speed, range_fuel, cost_efficiency) already uses. Two
builds of the same combo with wildly different actual masses scored
identically on cargo capacity, and the explore sliders had no effect on
it at all.

Moved the cargo_capacity/cargo_capacity_kg calculation into
_raw_physics_from_masses, deriving it from floor_total (the real
assembled mass) the same deadweight/lightship-ratio way as before, just
against the right mass. Removed the cargo_capacity_kg parameter that
threaded a precomputed constant through _decide_masses and
_raw_physics_from_masses -- it's now computed fresh at each point the
optimizer/explore sliders try, exactly like power_density/speed already
are. cost_efficiency's $/(kg·m) division now uses whichever cargo
convention the domain actually scores (cargo_capacity's 2.5x ratio or
cargo_capacity_kg's 0.3x ratio) instead of always assuming the former,
so the cost-per-cargo-kg number and the cargo capacity shown alongside
it always agree.

Confirmed on combo #876: cargo_capacity_kg moved from 5.7kg (declared
floor sum: 5+5+9=19kg * 0.3) to 18kg (actual 60kg build * 0.3), and now
responds to the explore sliders (9.4kg-35.2kg across actuator/storage
choices) the way every other metric already did.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 20:20:45 -05:00
3ed3918964 add speed as a derived metric, fix aerodynamic drag gap it exposed, expand urban_commuting
target_velocity was only ever a platform-declared input used to size the
actuator -- an achieved-speed OUTPUT never existed anywhere, even though
trip time clearly matters for a domain like urban commuting. Added
"speed" as a genuine derived metric: achieved steady-state cruise speed
computed from the build's own power_density and the medium's resistance,
the same way power_density/range_fuel/cost_efficiency are already
outputs of a build rather than inputs to it.

That immediately surfaced a known, previously-deferred gap: the
resistance model was mass-proportional only (rolling resistance), with
no velocity-squared aerodynamic drag term, so inverting power/resistance
for speed had no ceiling at all -- light vehicles were "achieving"
thousands of m/s. Added DRAG_POWER_COEFF_BY_MEDIUM (ground only, a
car-like reference cross-section) and a closed-form cubic solve
(_solve_achievable_speed_mps, via Cardano's formula, no iteration) for
the achieved speed where propulsive power balances resistance + drag.
Reused the same effective (drag-inclusive) resistance for range_fuel and
cost_efficiency's operating-cost term, since they're the same physical
quantity (energy spent per meter) evaluated at the build's actual speed.

This also closes the range-overestimation bug flagged much earlier
against combo #876 (a real e-bike): range dropped from ~1,032km to
~53km, right in the ~50-80km realistic e-bike range that was the
original target. Air and water media are unchanged (air's L/D-based
cruise model doesn't have this problem; water hull drag needs its own
treatment, not a car's frontal area -- left as a known remaining gap).

Also added cargo_capacity_kg to urban_commuting (whether a commute
vehicle can carry groceries/passengers/gear matters as much as the
metrics already scored there) and renormalized weights across the now
five metrics.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 19:59:45 -05:00
d1f14dbf14 remove food from ambient energy forms, size solar sails to actual power needs
Biological Feed (food) was treated as an "ambient" energy source, so
range_fuel always reported the domain's ceiling regardless of how much
food was carried -- stopping to eat is a resupply, the same category as
refuelling a tank, not a genuinely external/inexhaustible source like
sun or wind. Removed "biological" from AMBIENT_ENERGY_FORMS; food now
uses the normal storage-mass-limited range formula like any fuel.

Solar Sail was still special-cased to a fixed footprint-derived mass
(100m^2 -> 5kg) regardless of what a domain's power/velocity target
actually needed -- the same "fixed reference instead of a requirement
floor" bug biological actuators had before last session's fix. Folded
it into the same general requirement-floor + joint-optimizer path:
declared footprint becomes a FLOOR (SAIL_AREAL_DENSITY_KG_PER_M2), not
a fixed value, so sail size scales with what's actually needed --
"enough panels to supply enough power for actuator impulse" is now
enforced the same way structural/mass-ceiling requirements already are,
instead of relying on a product-spec constant that happened to work or
not. No actuator type is special-cased for mass sizing anymore.

Confirmed the fix surfaces an honest result rather than hiding one: a
solar sail's declared 0.01 W/kg specific power can never reach
interplanetary_travel's 10 W/kg power_density floor at any sail size
(the ratio is capped by the sail's own power_density regardless of
scale), so it correctly still scores 0 there -- a real technology/domain
mismatch, not a sizing bug.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 19:37:38 -05:00
6cdd308583 treat biological actuators as normal budget-competing mass, fix explore-panel visibility
Biological actuators (Human Muscle, Animal Traction) were special-cased
out of the mass optimizer entirely: a fixed 70kg reference used only in
the power formula, excluded from the platform's mass budget and from
the explore-panel sliders. That made "bigger operator" or "more
operators" inexpressible, and required a power_mass/denom_offset
parameter pair throughout the physics code solely to keep this one
case's numerator mass separate from its budget mass.

Operator mass is now a normal, budget-competing, structurally-carried
variable sized by the same joint optimizer as any mechanical actuator,
with BIOLOGICAL_OPERATOR_MASS_KG reinterpreted as a floor (at least one
real operator) rather than a fixed value -- the explore slider now
reads as "how many/how large are the operators." Since every remaining
case set power_mass == actuator_mass and denom_offset == 0.0 anyway,
those parameters were entirely vestigial once biological's special
case was gone, so _raw_physics_from_masses drops them.

Also: the explore section was gated on `explore_result is not none`,
so combos with no free mass to explore (radiation-pressure sails, or
previously biological) showed nothing at all instead of the existing
explanatory message. Gated on `scores` instead, so the section always
renders and the message inside `_explore_result.html` is reachable.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 19:17:59 -05:00
d871635779 score-optimize actuator/storage/platform allocation, enforce structural feasibility
The saved composite score previously came from a requirement-solve that
only satisfied the platform's physical performance floor, not the
domain's actual weighted score -- a smaller/cheaper build could always
score higher by hand. _decide_masses now jointly searches platform,
actuator, and storage mass (coarse-to-fine grid, no external deps) to
maximize the domain's real weighted composite score, with the
requirement floor as a lower bound rather than the final answer.

Platform mass specifically was previously fixed at a geometric-mean
representative value, which could be too little structure to carry its
own required actuator+storage (reusing CARGO_KG_PER_STRUCTURAL_KG, the
existing structure-carries-N-times-its-mass ratio, applied to a
platform carrying its own powertrain instead of cargo). Growing
platform mass also raises that structural ceiling, so it has to be
searched jointly rather than fixed or bounded independently.

Because power_density/range_fuel/cost_efficiency are all per-kg
ratios, none of them naturally penalize a build whose absolute mass
exceeds its own platform's declared ceiling -- a Piston Engine sized
for a Hyperloop could still score well on a Light Personal Vehicle.
Pass 2 now detects genuine infeasibility (no platform mass within its
own declared ceiling can structurally carry the required floor) and
saves it as a per-domain block instead of a misleadingly good score.

Also adds an explore-panel warning (not a hard block, since exploration
is intentionally loose) when a manually-dragged slider build exceeds
the platform's mass ceiling or structural carrying capacity.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 18:59:20 -05:00
76f460499a drop safety/availability from scoring, holistic p4 rating, phase-parallel pipeline
safety and availability don't reduce to physics formulas the way
power_density/range_fuel/cost_efficiency do -- they're judgment calls
(risk assessment, infrastructure prevalence), and running them through the
same log-normalize() built for physical quantities produced incoherent
results: safety's raw value is already a "0-1" score, and normalizing it
again turned 0.6 into an unexplainable 0.678 that even the LLM reviewing
it could only cite, never justify (see combo 1540). Removed both from
domain_metric_weights (safety from 4 domains, availability from
urban_commuting) and renormalized the remaining weights to sum to 1.0.

Pass 4 now produces one holistic RATING (LOW/MEDIUM/HIGH) alongside the
existing VERDICT, with safety and accessibility folded in as qualitative
considerations feeding that single judgment rather than scored
separately -- not a checklist of independent numbers. New
qualitative_rating column, filterable in the results UI. Also added
domain name/description to the review prompt so the LLM judges a metric
like range against what the domain actually needs (urban_commuting:
1-50km) instead of generic real-world expectations for the platform
category -- confirmed live on a combo where phi4 had called a 396km range
"limited" by comparing to typical aircraft rather than a domain that
needs 1-50km.

Pass 2 is estimator-only now -- self.llm is never consulted there,
reserved entirely for pass 4. Restructured Pipeline.run() from combo-first
to phase-parallel: each pass now runs to completion across every combo
before the next pass starts, rather than walking each combo through all
four passes before the next combo. This surfaced a real bug: domain-
blocked combos (status stays "valid" by design, not "_fail") were
slipping past a naive status-based skip guard and getting silently
re-processed by pass 2. Fixed with a shared dead-combo check that catches
both generic failures and domain blocks correctly.

Also fixes a results-page display bug found while reviewing a live combo:
the per-metric "position" bar showed raw distance from norm_min without
inverting for lower_is_better metrics, so an excellent cost score (near
the good end) rendered as a ~0%, near-empty bar -- looked bad next to its
own 0.99 normalized score.

Validated live against phi4 (real Ollama calls, not mocked): full-domain
phase-parallel run (2,970 combos, estimator-only p2, 1.6s) followed by a
real pass-4 run (111 reviewed, 11m, 0 crashes, 0 null ratings). Two tests
that relied on the old LLM-driven pass 2 to force deterministic outcomes
were updated to test pass 4's verdict-wiring directly instead. All 100
tests pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 16:52:26 -05:00
730a23bac3 give pass 4 raw values, batch deferrable commits, fix two known bugs
Pass 4's plausibility review only ever saw normalized scores, never the
raw physical estimate behind them -- confirmed via live testing this was
exactly what caused a real misfire (gemma2:27b cited a real cyclist's
correct 5 W/kg, log-normalized to "0.159" against a car's power scale, as
grounds for rejecting an ordinary bicycle). review_plausibility now takes
raw_metrics + normalized_scores + metric units, and the prompt explicitly
instructs reasoning from the raw value first. Verified live against phi4:
it now cites the actual raw number and correctly explains why a low
normalized score doesn't mean the estimate or concept is bad.

Repository write methods used in the pipeline's hot path now take an
optional commit=False, and Pipeline defers commits during the fast/
deterministic passes (1, 3, and 2 without an LLM), flushing every 200
combos and on any exit path (finally block covers normal completion,
cancellation, and any other exception). LLM-involving calls (pass 2 with
an LLM, all of pass 4) still commit immediately -- those are slow and
crash-prone and worth protecting per-write; the deterministic passes
aren't, and recomputing them is now measured at under a second for the
full domain rather than worth 8,000+ individual fsync'd commits. Full
2,970-combination domain run: multiple minutes -> 0.91s. Test suite:
~70s -> ~15s.

Also fixes two more issues found while auditing the estimator for a real
run: CARGO_KG_PER_STRUCTURAL_KG was 500 (no real vehicle carries 500x its
own structural mass in cargo -- a magnitude bug, not a modeling choice),
corrected to 2.5. And space/rocket platforms' range_fuel now reports the
domain's ceiling instead of an arbitrary placeholder constant -- vacuum
coast isn't resistance-limited, so "distance before running out of fuel"
isn't a meaningful question for these the way it is for ground/air/water
vehicles; the real constraint is delta-v budget, a different metric this
pass doesn't model.

Validated with a live full-domain run (phi4, real Ollama calls): 115
reviewed, 0 malformed/null reviews, 0 verdict-vs-status mismatches.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 15:23:45 -05:00
be25a837ff replace stub estimator with a requirement-derived physics engine
pass 2's no-LLM fallback previously used broken/placeholder formulas:
power_density passed an actuator's own intensive W/kg straight through
without scaling by vehicle mass, range_fuel multiplied energy_density by a
flat unitless constant, and cost_efficiency was a categorical guess.

Replaces all three with formulas grounded in each entity's own declared
attributes. Actuator and storage mass are sized to what's actually
necessary -- enough power to sustain a platform's target_velocity against
resistance (or real thrust/accel requirements where already declared, for
aircraft/rocket combos), enough energy to reach the domain's own declared
range ceiling -- solved as a closed-form 2x2 linear system rather than an
invented mass-fraction table. Platform mass uses the geometric mean of its
declared range instead of the bare floor, since a category as broad as
Road Vehicle (50kg-36,000kg) is closer to log-uniformly distributed than
uniformly distributed. cost_efficiency is now real operating cost (energy
price x resistance) plus amortized upfront cost (materials cost by medium
x lifetime distance), replacing the old flat per-energy-form guess.

Also fixes two bugs found while validating the above against real-world
reference values: entities whose power source isn't their own carried mass
(Human Muscle, Solar Sail) degenerated to zero power; and combos blocked
by a domain-specific constraint kept combinations.status stuck at "valid"
forever, which silently miscounted them as passing in two repository
queries (count_combinations_by_status, get_pipeline_summary) even though
the per-domain result row was correctly marked blocked.

Adds target_velocity to platforms that had no declared performance
requirement at all, and raises OllamaLLMProvider's HTTP timeout (120s to
300s) to match observed real review-call latency.

Validated against 4 real-world reference combos (commuter car, bicycle,
delivery drone) and a full 2,970-combination domain run (0 pass-2
failures); all 100 existing tests still pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 13:45:01 -05:00
33 changed files with 2984 additions and 494 deletions

View File

@@ -60,7 +60,7 @@ tests/ # pytest, uses seeded_repo fixture from conftest.py
## Data flow (pipeline passes)
1. **Pass 1 — Constraints**: `ConstraintResolver.resolve()` → blocked/conditional/valid. Blocked combos get a result row and `continue`.
2. **Pass 2 — Estimation**: LLM or `_stub_estimate()` → raw metric values. Saved immediately via `save_raw_estimates()` (normalized_score=NULL).
2. **Pass 2 — Estimation**: `_estimate_physics()` (deterministic physics engine; estimator-only, no LLM) → raw metric values. Saved immediately via `save_raw_estimates()` (normalized_score=NULL).
3. **Pass 3 — Scoring**: `Scorer.score_combination()` → log-normalized scores + weighted geometric mean composite. Saves via `save_scores()` + `save_result()`.
4. **Pass 4 — LLM Review**: Only for above-threshold combos with an LLM provider. No real provider yet (only `MockLLMProvider`).
5. **Pass 5 — Human Review**: Manual via web UI results page.

View File

@@ -0,0 +1,34 @@
# GPU batching for the mass-allocation optimizer — not yet, here's the threshold
## Context
`Pipeline._decide_masses`'s joint platform/actuator/storage optimizer (coarse-to-fine grid search, see `_search_best_allocation`) calls its objective function roughly 11,700 times per combo. Profiling confirmed this dominates pipeline runtime: for 50 combos, 582,920 objective-function calls, each doing scalar arithmetic (power_density, the drag cubic solve, normalize, composite_score) on one `(platform, actuator, storage)` triple. The cost is Python's per-call overhead (bytecode dispatch, refcounting, attribute lookups), not the arithmetic itself — the individual formulas are cheap.
## Why GPU doesn't fit today
A single combo's grid is only ~169 points per round (13×13). GPUs pay off when there's enough independent parallel work to amortize kernel-launch and host↔device transfer overhead (each typically tens of microseconds to low milliseconds); 169 elements doesn't come close, and that overhead would be paid repeatedly — once per grid round, ~6-10 rounds per combo.
The parallelism that actually exists is **across combos**, not within one combo's grid — every combo's optimization is fully independent of every other's. At current scale (~180 combos reach the optimizer per domain after Pass 1 filtering), batching every combo's grid into one array gives ~180×169 ≈ 30K elements per round — borderline, probably a wash against plain CPU numpy.
## The actual threshold
Combo count scales **multiplicatively** with added dimensions or entities per dimension (today: 11 platforms × 15 actuators × 18 storages ≈ 2,970 combos, ~180 of which reach the optimizer). Add a 4th dimension with even 10 options and total combos scale to ~30,000, with optimizer-eligible combos likely growing roughly proportionally to ~1,800/domain — batched grid size ≈ 300K elements/round. A 5th dimension does it again, into the low millions. That's the regime where a GPU's thousands of cores start meaningfully outrunning a CPU's 4-16-wide SIMD lanes.
So: not "more dimensions" directly, but the combo×grid batch size those dimensions produce. Rough rule of thumb from this discussion:
- **Tens of thousands of elements/round** (current scale, or a modest one-dimension addition): plain CPU numpy vectorization is enough, no GPU.
- **Hundreds of thousands to low millions**: GPU batching across combos starts being worth evaluating.
## What GPU batching would actually require
Not just "swap numpy for cupy." It means restructuring `_process_pass2` from combo-first (one combo through the optimizer at a time) to batch-first (a chunk of N combos' grids evaluated together as one array with a "combo" axis, broadcasting each combo's own constants — `k_act`, `k_med`, `e_dens`, drag coefficients, mass bounds — across that axis). That's a real architectural change, not a drop-in acceleration:
- **CLAUDE.md documents the pipeline as deliberately combo-first**: "each combo goes through all requested passes before the next combo starts... Progress is persisted per-combo (crash-safe, resumable)." Batching means checkpointing per-*batch*, not per-combo — a real (if manageable) tradeoff against that resumability guarantee, not something that comes free alongside the speedup.
- Every branch in `_raw_physics_from_masses` / `_solve_achievable_speed_mps` (biological floors, ambient energy forms, degenerate fallbacks, the cubic's edge cases) needs to become `np.where(condition, a, b)` instead of `if/else` — careful, error-prone translation work, not mechanical.
## Decision
Don't build this now — no current need at ~180 combos/domain. If dimension count grows enough to matter, do it in two steps:
1. **CPU numpy vectorization first** (batch one combo's grid into arrays, evaluate with vectorized ops instead of a Python double-loop). This is needed regardless of GPU or not, since it's the same rewrite either way, and profiling suggests it could plausibly give 10-50x on its own by replacing ~11,700 Python calls/combo with a couple dozen numpy batch calls.
2. **Re-profile at the new scale.** Only reach for GPU batching-across-combos if CPU numpy is still the dominant cost after step 1, and only once the batched element count is actually in GPU-favorable territory (see thresholds above) — this is a "measure, then decide" call, not something to build ahead of need.

View File

@@ -9,7 +9,7 @@ from datetime import datetime, timezone
from typing import Sequence
from physcom.models.entity import Dependency, Entity
from physcom.models.domain import Domain, DomainConstraint, MetricBound
from physcom.models.domain import Domain, DomainConstraint, FreeVariable, MetricBound, MetricFormula
from physcom.models.combination import Combination
@@ -20,6 +20,14 @@ class Repository:
self.conn = conn
self.conn.row_factory = sqlite3.Row
def commit(self) -> None:
"""Explicit flush point, for callers batching writes with commit=False
below (see Pipeline.run: instant/deterministic passes defer commits
and flush in bulk, since a crash there just means cheap recompute;
LLM-call results still commit immediately, since those are slow/
expensive to redo)."""
self.conn.commit()
# ── Dimensions ──────────────────────────────────────────────
def ensure_dimension(self, name: str, description: str = "") -> int:
@@ -222,24 +230,37 @@ class Repository:
self.conn.commit()
return row["id"]
def backfill_lower_is_better(self, domain_name: str, metric_name: str) -> None:
"""Set lower_is_better=1 for an existing domain-metric row that still has the default 0."""
def sync_domain_metric_weights(self, domain: Domain) -> None:
"""Make domain_metric_weights exactly match domain.metric_bounds on an
already-seeded domain: upserts weight/norm_min/norm_max/unit for every
currently-declared metric, and deletes any row for a metric that's been
removed from the domain (e.g. safety/availability dropped from the
scored set). Safe to call whether the domain was just freshly inserted
or already existed.
"""
row = self.conn.execute(
"SELECT id FROM domains WHERE name = ?", (domain.name,)
).fetchone()
if not row:
return
domain_id = row["id"]
keep_ids = []
for mb in domain.metric_bounds:
metric_id = self.ensure_metric(mb.metric_name, unit=mb.unit)
keep_ids.append(metric_id)
self.conn.execute(
"""UPDATE domain_metric_weights SET lower_is_better = 1
WHERE lower_is_better = 0
AND domain_id = (SELECT id FROM domains WHERE name = ?)
AND metric_id = (SELECT id FROM metrics WHERE name = ?)""",
(domain_name, metric_name),
"""INSERT OR REPLACE INTO domain_metric_weights
(domain_id, metric_id, weight, norm_min, norm_max, lower_is_better, unit)
VALUES (?, ?, ?, ?, ?, ?, ?)""",
(domain_id, metric_id, mb.weight, mb.norm_min, mb.norm_max,
int(mb.lower_is_better), mb.unit),
)
self.conn.commit()
def backfill_metric_unit(self, domain_name: str, metric_name: str, unit: str) -> None:
"""Set this domain-metric row's unit — unit is domain-scoped, not global to the metric name."""
if keep_ids:
placeholders = ",".join("?" * len(keep_ids))
self.conn.execute(
"""UPDATE domain_metric_weights SET unit = ?
WHERE domain_id = (SELECT id FROM domains WHERE name = ?)
AND metric_id = (SELECT id FROM metrics WHERE name = ?)""",
(unit, domain_name, metric_name),
f"""DELETE FROM domain_metric_weights
WHERE domain_id = ? AND metric_id NOT IN ({placeholders})""",
(domain_id, *keep_ids),
)
self.conn.commit()
@@ -265,6 +286,10 @@ class Repository:
"INSERT OR IGNORE INTO domain_constraints (domain_id, key, value) VALUES (?, ?, ?)",
(domain.id, dc.key, val),
)
for fv in domain.free_variables:
self.add_free_variable(domain.id, fv, commit=False)
for mf in domain.metric_formulas:
self.add_metric_formula(domain.id, mf, commit=False)
self.conn.commit()
return domain
@@ -278,6 +303,31 @@ class Repository:
by_key.setdefault(r["key"], []).append(r["value"])
return [DomainConstraint(key=k, allowed_values=v) for k, v in by_key.items()]
def _load_free_variables(self, domain_id: int) -> list[FreeVariable]:
rows = self.conn.execute(
"""SELECT id, name, sort_order, floor_formula, ceiling_formula
FROM domain_free_variables WHERE domain_id = ? ORDER BY sort_order""",
(domain_id,),
).fetchall()
return [
FreeVariable(
id=r["id"], name=r["name"], sort_order=r["sort_order"],
floor_formula=r["floor_formula"], ceiling_formula=r["ceiling_formula"],
)
for r in rows
]
def _load_metric_formulas(self, domain_id: int) -> list[MetricFormula]:
rows = self.conn.execute(
"""SELECT id, metric_name, formula
FROM domain_metric_formulas WHERE domain_id = ? ORDER BY metric_name""",
(domain_id,),
).fetchall()
return [
MetricFormula(id=r["id"], metric_name=r["metric_name"], formula=r["formula"])
for r in rows
]
def _load_domain(self, where: str, param: str | int) -> Domain | None:
row = self.conn.execute(f"SELECT * FROM domains WHERE {where} = ?", (param,)).fetchone()
if not row:
@@ -304,6 +354,8 @@ class Repository:
for w in weights
],
constraints=self._load_domain_constraints(row["id"]),
free_variables=self._load_free_variables(row["id"]),
metric_formulas=self._load_metric_formulas(row["id"]),
)
def get_domain(self, name: str) -> Domain | None:
@@ -355,12 +407,63 @@ class Repository:
)
self.conn.commit()
# ── Free variables & metric formulas ──────────────────────────
def add_free_variable(self, domain_id: int, fv: FreeVariable, commit: bool = True) -> FreeVariable:
cur = self.conn.execute(
"""INSERT INTO domain_free_variables
(domain_id, name, sort_order, floor_formula, ceiling_formula)
VALUES (?, ?, ?, ?, ?)""",
(domain_id, fv.name, fv.sort_order, fv.floor_formula, fv.ceiling_formula),
)
fv.id = cur.lastrowid
if commit:
self.conn.commit()
return fv
def update_free_variable(self, fv_id: int, fv: FreeVariable) -> None:
self.conn.execute(
"""UPDATE domain_free_variables
SET name = ?, sort_order = ?, floor_formula = ?, ceiling_formula = ?
WHERE id = ?""",
(fv.name, fv.sort_order, fv.floor_formula, fv.ceiling_formula, fv_id),
)
self.conn.commit()
def delete_free_variable(self, fv_id: int) -> None:
self.conn.execute("DELETE FROM domain_free_variables WHERE id = ?", (fv_id,))
self.conn.commit()
def add_metric_formula(self, domain_id: int, mf: MetricFormula, commit: bool = True) -> MetricFormula:
cur = self.conn.execute(
"""INSERT OR REPLACE INTO domain_metric_formulas (domain_id, metric_name, formula)
VALUES (?, ?, ?)""",
(domain_id, mf.metric_name, mf.formula),
)
mf.id = cur.lastrowid
if commit:
self.conn.commit()
return mf
def update_metric_formula(self, mf_id: int, mf: MetricFormula) -> None:
self.conn.execute(
"UPDATE domain_metric_formulas SET metric_name = ?, formula = ? WHERE id = ?",
(mf.metric_name, mf.formula, mf_id),
)
self.conn.commit()
def delete_metric_formula(self, mf_id: int) -> None:
self.conn.execute("DELETE FROM domain_metric_formulas WHERE id = ?", (mf_id,))
self.conn.commit()
def delete_domain(self, domain_id: int) -> None:
self.conn.execute("DELETE FROM pipeline_runs WHERE domain_id = ?", (domain_id,))
self.conn.execute("DELETE FROM combination_results WHERE domain_id = ?", (domain_id,))
self.conn.execute("DELETE FROM combination_scores WHERE domain_id = ?", (domain_id,))
self.conn.execute("DELETE FROM domain_metric_weights WHERE domain_id = ?", (domain_id,))
self.conn.execute("DELETE FROM domain_constraints WHERE domain_id = ?", (domain_id,))
self.conn.execute("DELETE FROM domain_free_variables WHERE domain_id = ?", (domain_id,))
self.conn.execute("DELETE FROM domain_metric_formulas WHERE domain_id = ?", (domain_id,))
self.conn.execute("DELETE FROM domains WHERE id = ?", (domain_id,))
self.conn.commit()
@@ -422,7 +525,7 @@ class Repository:
key = ",".join(str(eid) for eid in sorted(entity_ids))
return hashlib.sha256(key.encode()).hexdigest()[:16]
def save_combination(self, combination: Combination) -> Combination:
def save_combination(self, combination: Combination, commit: bool = True) -> Combination:
entity_ids = [e.id for e in combination.entities]
combination.hash = self.compute_hash(entity_ids)
@@ -446,14 +549,15 @@ class Repository:
"INSERT INTO combination_entities (combination_id, entity_id) VALUES (?, ?)",
(combination.id, eid),
)
if commit:
self.conn.commit()
return combination
def update_combination_status(
self, combo_id: int, status: str, block_reason: str | None = None
self, combo_id: int, status: str, block_reason: str | None = None, commit: bool = True
) -> None:
# Don't downgrade from higher pass states — preserves human/LLM review data
if status in ("scored", "llm_reviewed") or status.endswith("_fail"):
if status in ("scored", "llm_reviewed", "valid") or status.endswith("_fail"):
row = self.conn.execute(
"SELECT status FROM combinations WHERE id = ?", (combo_id,)
).fetchone()
@@ -466,10 +570,18 @@ class Repository:
return
if status == "llm_reviewed" and cur == "reviewed":
return
# "valid" is pass 1's domain-agnostic result -- a combo
# already at any later pass state (or a fail state) has
# progressed past pass 1 already, in this domain or
# another one sharing the same combo. Pass 1 re-running
# for a different domain must not silently revert that.
if status == "valid" and cur not in (None, "valid"):
return
self.conn.execute(
"UPDATE combinations SET status = ?, block_reason = ? WHERE id = ?",
(status, block_reason, combo_id),
)
if commit:
self.conn.commit()
def get_combination(self, combo_id: int) -> Combination | None:
@@ -558,6 +670,7 @@ class Repository:
combo_id: int,
domain_id: int,
scores: list[dict],
commit: bool = True,
) -> None:
"""Save per-metric scores. Each dict: metric_id, raw_value, normalized_score, estimation_method, confidence."""
for s in scores:
@@ -569,6 +682,7 @@ class Repository:
(combo_id, domain_id, s["metric_id"], s["raw_value"],
s["normalized_score"], s["estimation_method"], s["confidence"]),
)
if commit:
self.conn.commit()
def save_result(
@@ -581,15 +695,20 @@ class Repository:
llm_review: str | None = None,
human_notes: str | None = None,
domain_block_reason: str | None = None,
qualitative_rating: str | None = None,
commit: bool = True,
) -> None:
self.conn.execute(
"""INSERT OR REPLACE INTO combination_results
(combination_id, domain_id, composite_score, novelty_flag,
llm_review, human_notes, pass_reached, domain_block_reason)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)""",
llm_review, human_notes, pass_reached, domain_block_reason,
qualitative_rating)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)""",
(combo_id, domain_id, composite_score, novelty_flag,
llm_review, human_notes, pass_reached, domain_block_reason),
llm_review, human_notes, pass_reached, domain_block_reason,
qualitative_rating),
)
if commit:
self.conn.commit()
def get_combination_scores(self, combo_id: int, domain_id: int) -> list[dict]:
@@ -606,13 +725,19 @@ class Repository:
def count_combinations_by_status(self, domain_name: str | None = None) -> dict[str, int]:
"""Count combos by status. If domain_name given, only combos with results in that domain."""
if domain_name:
# combinations.status is domain-agnostic (a combo can be "valid"
# generically but blocked by one domain's own constraints), so a
# domain-scoped count must bucket domain_block_reason rows on
# their own rather than trusting c.status.
rows = self.conn.execute(
"""SELECT c.status, COUNT(*) as cnt
"""SELECT CASE WHEN cr.domain_block_reason IS NOT NULL
THEN 'domain_blocked' ELSE c.status END as status,
COUNT(*) as cnt
FROM combination_results cr
JOIN combinations c ON cr.combination_id = c.id
JOIN domains d ON cr.domain_id = d.id
WHERE d.name = ?
GROUP BY c.status""",
GROUP BY status""",
(domain_name,),
).fetchall()
else:
@@ -621,6 +746,20 @@ class Repository:
).fetchall()
return {r["status"]: r["cnt"] for r in rows}
def count_results_by_rating(self, domain_name: str) -> dict[str, int]:
"""Count results by qualitative_rating (LOW/MEDIUM/HIGH) for a domain.
Rows with no rating (not yet pass-4 reviewed, or reviewed before this
existed) are excluded, not bucketed as a pseudo-status."""
rows = self.conn.execute(
"""SELECT cr.qualitative_rating as rating, COUNT(*) as cnt
FROM combination_results cr
JOIN domains d ON cr.domain_id = d.id
WHERE d.name = ? AND cr.qualitative_rating IS NOT NULL
GROUP BY cr.qualitative_rating""",
(domain_name,),
).fetchall()
return {r["rating"]: r["cnt"] for r in rows}
def get_pipeline_summary(self, domain_name: str) -> dict | None:
"""Return a summary of results for a domain, or None if no results."""
row = self.conn.execute(
@@ -641,7 +780,8 @@ class Repository:
FROM combinations c
JOIN combination_results cr ON cr.combination_id = c.id
JOIN domains d ON cr.domain_id = d.id
WHERE c.status LIKE '%\\_fail' ESCAPE '\\' AND d.name = ?""",
WHERE (c.status LIKE '%\\_fail' ESCAPE '\\' OR cr.domain_block_reason IS NOT NULL)
AND d.name = ?""",
(domain_name,),
).fetchone()
return {
@@ -664,17 +804,26 @@ class Repository:
).fetchone()
return dict(row) if row else None
def get_all_results(self, domain_name: str, status: str | None = None) -> list[dict]:
"""Return all results for a domain, optionally filtered by combo status."""
def get_all_results(
self, domain_name: str, status: str | None = None, rating: str | None = None
) -> list[dict]:
"""Return all results for a domain, optionally filtered by combo
status and/or qualitative_rating (LOW/MEDIUM/HIGH, independent filters
that combine with AND)."""
query = """SELECT cr.*, c.hash, c.status as combo_status, d.name as domain_name
FROM combination_results cr
JOIN combinations c ON cr.combination_id = c.id
JOIN domains d ON cr.domain_id = d.id
WHERE d.name = ?"""
params: list = [domain_name]
if status:
query += " AND c.status = ?"
if status == "domain_blocked":
query += " AND cr.domain_block_reason IS NOT NULL"
elif status:
query += " AND c.status = ? AND cr.domain_block_reason IS NULL"
params.append(status)
if rating:
query += " AND cr.qualitative_rating = ?"
params.append(rating)
query += " ORDER BY cr.composite_score DESC"
rows = self.conn.execute(query, params).fetchall()
combo_ids = [r["combination_id"] for r in rows]
@@ -689,6 +838,7 @@ class Repository:
"pass_reached": r["pass_reached"],
"domain_id": r["domain_id"],
"domain_block_reason": r["domain_block_reason"],
"qualitative_rating": r["qualitative_rating"],
}
for r in rows
]
@@ -798,7 +948,7 @@ class Repository:
return row["pass_reached"] if row else None
def save_raw_estimates(
self, combo_id: int, domain_id: int, estimates: list[dict]
self, combo_id: int, domain_id: int, estimates: list[dict], commit: bool = True
) -> None:
"""Save raw metric estimates (pass 2) with normalized_score=NULL.
@@ -813,6 +963,7 @@ class Repository:
(combo_id, domain_id, e["metric_id"], e["raw_value"],
e["estimation_method"], e["confidence"]),
)
if commit:
self.conn.commit()
def get_existing_result(self, combo_id: int, domain_id: int) -> dict | None:
@@ -837,6 +988,8 @@ class Repository:
self.conn.execute("DELETE FROM entities")
self.conn.execute("DELETE FROM domain_metric_weights")
self.conn.execute("DELETE FROM domain_constraints")
self.conn.execute("DELETE FROM domain_free_variables")
self.conn.execute("DELETE FROM domain_metric_formulas")
self.conn.execute("DELETE FROM domains")
self.conn.execute("DELETE FROM metrics")
self.conn.execute("DELETE FROM dimensions")

View File

@@ -91,6 +91,7 @@ CREATE TABLE IF NOT EXISTS combination_results (
human_notes TEXT,
pass_reached INTEGER,
domain_block_reason TEXT,
qualitative_rating TEXT,
UNIQUE(combination_id, domain_id)
);
@@ -119,6 +120,24 @@ CREATE TABLE IF NOT EXISTS domain_constraints (
UNIQUE(domain_id, key, value)
);
CREATE TABLE IF NOT EXISTS domain_free_variables (
id INTEGER PRIMARY KEY AUTOINCREMENT,
domain_id INTEGER NOT NULL REFERENCES domains(id),
name TEXT NOT NULL,
sort_order INTEGER NOT NULL,
floor_formula TEXT NOT NULL,
ceiling_formula TEXT NOT NULL,
UNIQUE(domain_id, name)
);
CREATE TABLE IF NOT EXISTS domain_metric_formulas (
id INTEGER PRIMARY KEY AUTOINCREMENT,
domain_id INTEGER NOT NULL REFERENCES domains(id),
metric_name TEXT NOT NULL,
formula TEXT NOT NULL,
UNIQUE(domain_id, metric_name)
);
CREATE INDEX IF NOT EXISTS idx_deps_entity ON dependencies(entity_id);
CREATE INDEX IF NOT EXISTS idx_deps_category_key ON dependencies(category, key);
CREATE INDEX IF NOT EXISTS idx_combo_status ON combinations(status);
@@ -165,6 +184,10 @@ def _migrate(conn: sqlite3.Connection) -> None:
conn.execute(
"ALTER TABLE combination_results ADD COLUMN domain_block_reason TEXT"
)
if "qualitative_rating" not in result_cols:
conn.execute(
"ALTER TABLE combination_results ADD COLUMN qualitative_rating TEXT"
)
# Backfill: cost_efficiency is lower-is-better in all domains
conn.execute(

View File

@@ -32,6 +32,13 @@ CATEGORY_SEVERITY: dict[str, str] = {
"energy": "block",
"environment": "block",
"infrastructure": "skip",
# Safety-critical physical necessities (radiation shielding, containment,
# etc.) -- same severity as energy/environment, not the softer default
# "warn" every other category falls through to. Missing this entry meant
# Nuclear Thermal Drive/Nuclear Fuel's "material" requires (radiation_
# shielding) defaulted to a non-blocking warning nothing in the catalog
# ever satisfies -- see LOGIC DOCS/002's "silent guardrail hole" pattern.
"material": "block",
}
# For provides-vs-range_min: deficit > this ratio = hard block, else warning
@@ -53,6 +60,40 @@ KEY_AGGREGATION: dict[str, str] = {
OVERRUN_TOLERANCE: float = 0.10
def aggregate_dependency_value(
combination: Combination,
key: str,
constraint_type: str,
key_aggregation: dict[str, str] | None = None,
) -> float | None:
"""Collapse every numeric dependency matching (key, constraint_type)
across a combination's entities into one system-level number: summed for
extensive keys (KEY_AGGREGATION says "sum", e.g. mass/footprint --
independent components sharing one physical vehicle), otherwise the
strongest single value wins (today's default pairwise behavior). Returns
None if no entity declares a matching numeric dependency. Shared by
ConstraintResolver._check_provides_vs_range and the formula evaluator's
injected dep() builtin (see engine/formula.py, engine/pipeline.py) so
there's one implementation of "how do these entities' declared numbers
combine," not two.
"""
key_aggregation = KEY_AGGREGATION if key_aggregation is None else key_aggregation
values: list[float] = []
for entity in combination.entities:
for dep in entity.dependencies:
if dep.key != key or dep.constraint_type != constraint_type:
continue
try:
values.append(float(dep.value))
except (ValueError, TypeError):
continue
if not values:
return None
if key_aggregation.get(key) == "sum":
return sum(values)
return max(values)
@dataclass
class ConstraintResult:
"""Outcome of constraint resolution for a combination."""
@@ -243,11 +284,13 @@ class ConstraintResolver:
required.setdefault(dep.key, []).append((entity.name, val))
for key in set(provided) & set(required):
prov_val = aggregate_dependency_value(
combination, key, "provides", self.key_aggregation
)
if self.key_aggregation.get(key) == "sum":
prov_name = " + ".join(name for name, _ in provided[key])
prov_val = sum(val for _, val in provided[key])
else:
prov_name, prov_val = max(provided[key], key=lambda t: t[1])
prov_name, _ = max(provided[key], key=lambda t: t[1])
for req_name, req_val in required[key]:
if prov_val < req_val * self.deficit_threshold:

View File

@@ -0,0 +1,139 @@
"""Safe arithmetic expression language for domain-authored estimator formulas.
No eval()/exec() anywhere -- compile_formula validates every AST node against
a fixed whitelist (arithmetic, numeric/string constants, name lookups, calls
to an explicitly supplied function table) before evaluate_formula ever walks
it, so a formula can express "how much drawback_force does this combo
produce" but never anything with attribute/subscript/import access.
"""
from __future__ import annotations
import ast
import math
import operator
from dataclasses import dataclass
from typing import Callable
class FormulaError(Exception):
"""Raised for invalid formula syntax/structure or a failed evaluation."""
_ALLOWED_NODES = (
ast.Expression, ast.BinOp, ast.UnaryOp, ast.Constant, ast.Name, ast.Load,
ast.Call, ast.keyword,
ast.Add, ast.Sub, ast.Mult, ast.Div, ast.Pow, ast.USub, ast.UAdd,
)
_BINOPS: dict[type, Callable[[float, float], float]] = {
ast.Add: operator.add,
ast.Sub: operator.sub,
ast.Mult: operator.mul,
ast.Div: operator.truediv,
ast.Pow: operator.pow,
}
_UNARYOPS: dict[type, Callable[[float], float]] = {
ast.USub: operator.neg,
ast.UAdd: operator.pos,
}
DEFAULT_FUNCTIONS: dict[str, Callable[..., float]] = {
"min": min,
"max": max,
"abs": abs,
"sqrt": math.sqrt,
"log": math.log,
"log1p": math.log1p,
"exp": math.exp,
}
@dataclass(frozen=True)
class CompiledFormula:
source: str
_tree: ast.Expression
def _validate(tree: ast.AST) -> None:
for node in ast.walk(tree):
if not isinstance(node, _ALLOWED_NODES):
raise FormulaError(
f"disallowed expression element: {type(node).__name__}"
)
if isinstance(node, ast.Constant):
if isinstance(node.value, bool) or not isinstance(node.value, (int, float, str)):
raise FormulaError(
f"disallowed constant type: {type(node.value).__name__}"
)
if isinstance(node, ast.Name) and node.id.startswith("__"):
raise FormulaError(f"disallowed name: {node.id}")
if isinstance(node, ast.Call) and not isinstance(node.func, ast.Name):
raise FormulaError("only direct function calls are allowed")
def compile_formula(source: str) -> CompiledFormula:
try:
tree = ast.parse(source, mode="eval")
except SyntaxError as exc:
raise FormulaError(f"invalid syntax in '{source}': {exc}") from exc
_validate(tree)
return CompiledFormula(source=source, _tree=tree)
def _eval(node: ast.AST, variables: dict[str, float], functions: dict[str, Callable]):
if isinstance(node, ast.Expression):
return _eval(node.body, variables, functions)
if isinstance(node, ast.Constant):
# Numeric constants are cast to float (not left as int) so a formula
# like a**b**c can't build an arbitrary-precision giant int before
# ever raising -- float exponentiation overflows to inf/OverflowError
# quickly instead. String constants (dep() key/constraint_type args)
# pass through unchanged.
return node.value if isinstance(node.value, str) else float(node.value)
if isinstance(node, ast.Name):
if node.id not in variables:
raise FormulaError(f"unknown variable '{node.id}'")
return variables[node.id]
if isinstance(node, ast.BinOp):
op = _BINOPS.get(type(node.op))
if op is None:
raise FormulaError(f"unsupported operator: {type(node.op).__name__}")
return op(
_eval(node.left, variables, functions),
_eval(node.right, variables, functions),
)
if isinstance(node, ast.UnaryOp):
op = _UNARYOPS.get(type(node.op))
if op is None:
raise FormulaError(f"unsupported operator: {type(node.op).__name__}")
return op(_eval(node.operand, variables, functions))
if isinstance(node, ast.Call):
fname = node.func.id # validated as ast.Name by _validate
func = functions.get(fname)
if func is None:
raise FormulaError(f"unknown function '{fname}'")
args = [_eval(a, variables, functions) for a in node.args]
kwargs = {kw.arg: _eval(kw.value, variables, functions) for kw in node.keywords}
return func(*args, **kwargs)
raise FormulaError(f"unsupported expression element: {type(node).__name__}")
def evaluate_formula(
compiled: CompiledFormula,
variables: dict[str, float],
functions: dict[str, Callable] | None = None,
) -> float:
effective_functions = {**DEFAULT_FUNCTIONS, **(functions or {})}
try:
result = _eval(compiled._tree, variables, effective_functions)
except FormulaError:
raise
except (TypeError, ValueError, ArithmeticError) as exc:
raise FormulaError(f"error evaluating '{compiled.source}': {exc}") from exc
if not isinstance(result, (int, float)) or isinstance(result, bool):
raise FormulaError(
f"formula '{compiled.source}' did not evaluate to a number"
)
return float(result)

File diff suppressed because it is too large Load Diff

View File

@@ -4,7 +4,7 @@ from __future__ import annotations
from abc import ABC, abstractmethod
from physcom.models.domain import MetricBound
from physcom.models.domain import Domain, MetricBound
class LLMRateLimitError(Exception):
@@ -36,8 +36,25 @@ class LLMProvider(ABC):
@abstractmethod
def review_plausibility(
self, combination_description: str, scores: dict[str, float]
self,
combination_description: str,
raw_metrics: dict[str, float],
normalized_scores: dict[str, float],
domain: Domain,
) -> tuple[str, bool]:
"""Given a combination and its scores, return a (text, is_plausible)
tuple: natural-language assessment and whether the concept is plausible."""
"""Given a combination, its raw physical estimates, and their
normalized scores, return a (text, is_plausible) tuple:
natural-language assessment and whether the concept is plausible.
Both raw_metrics and normalized_scores are given (not just the
normalized score) so the review can reason from the actual physics
rather than only a compressed 0-1 number, which can look
deceptively bad for a metric whose scale was built for a different
kind of vehicle. `domain` carries both each metric's unit (via
domain.metric_bounds, for formatting the raw value meaningfully)
and the domain's own name/description, so the review judges a
metric like range against what THIS domain actually needs rather
than generic real-world expectations for the platform category
(e.g. a short-hop domain shouldn't get judged against typical
long-haul aircraft range)."""
...

View File

@@ -16,6 +16,13 @@ def parse_verdict(text: str) -> bool:
return True
def parse_rating(text: str) -> str | None:
"""Extract RATING: LOW/MEDIUM/HIGH from response; None if absent (older
reviews saved before this existed, or a malformed response)."""
m = re.search(r"RATING:\s*(LOW|MEDIUM|HIGH)", text, re.IGNORECASE)
return m.group(1).upper() if m else None
def parse_metric_json(text: str, metrics: list[MetricBound]) -> dict[str, float]:
"""Strip markdown fences and parse JSON; fall back to each metric's own
norm_min/norm_max midpoint on error — a flat constant like 0.5 is

View File

@@ -20,6 +20,35 @@ def format_metrics_for_prompt(metrics: list["MetricBound"]) -> str:
return "\n".join(lines)
def format_scores_for_prompt(
raw_metrics: dict[str, float],
normalized_scores: dict[str, float],
metrics: list["MetricBound"],
) -> str:
"""Render each metric with BOTH its raw physical value and its
normalized score, so the reviewing pass can reason from the actual
physics instead of only ever seeing a compressed 0-1 number.
A real, correct estimate can still look damning once log-normalized
against a scale built for a different kind of vehicle (a cyclist's
real ~5 W/kg reads as "0.159" next to a car's 2000 W/kg ceiling) --
a reviewer that only sees the 0.159 has no way to notice that. See
the labeled-set calibration note on PLAUSIBILITY_REVIEW_PROMPT below.
"""
lines = []
for mb in metrics:
normed = normalized_scores.get(mb.metric_name)
if normed is None:
continue
raw = raw_metrics.get(mb.metric_name)
unit = mb.unit or "dimensionless"
raw_str = f"{raw:g} {unit}" if raw is not None else "unknown"
lines.append(
f"- {mb.metric_name}: raw estimate {raw_str} — normalized score {normed:.3f}"
)
return "\n".join(lines)
PHYSICS_ESTIMATION_PROMPT = """\
You are a physics estimation assistant. Given the following transportation concept, \
estimate the requested metrics using order-of-magnitude physics reasoning.
@@ -44,15 +73,22 @@ match that magnitude, don't guess a generically "reasonable-looking" decimal.
{{"some_metric": <number>, "another_metric": <number>}} — no explanatory text.
"""
# ponytail: pass 4 only sees pass 2's raw numbers, not its reasoning. Sharpened
# prompts on both sides closed most of the gap (a bad safety estimate went from
# 0.95 to 0.80 on the same combo once pass 2 was told to consider combination-
# specific hazards). Upgrade path if this isn't good enough in practice: have
# estimate_physics() also return a short per-metric reason, persist it
# alongside raw_value (new nullable column), and feed it into this prompt so
# pass 4 has something concrete to agree or disagree with. Deferred because it
# needs a schema/interface change across LLMProvider + both providers +
# pipeline + scorer + repository, and more generated tokens per combo.
# ponytail: pass 4 used to see only pass 2's normalized scores, not the raw
# physical numbers or any reasoning behind them. Fixed the raw-value half of
# that gap: format_scores_for_prompt() now shows both, since a correct raw
# estimate can look damning once log-normalized against a scale built for a
# different kind of vehicle (a cyclist's real ~5 W/kg reads as "0.159" next
# to a car's 2000 W/kg ceiling) -- gemma2:27b did exactly this on a real
# bicycle combo, citing "extremely low power density (0.159)" as grounds for
# IMPLAUSIBLE while never reasoning from the actual (correct) 5 W/kg. The
# reasoning-text half of the gap is still open: estimate_physics() doesn't
# return a per-metric rationale, so pass 4 still can't see WHY pass 2 landed
# on a number, only what the number is. Upgrade path if the raw value alone
# isn't enough in practice: have estimate_physics() also return a short
# per-metric reason, persist it alongside raw_value (new nullable column),
# and feed it into this prompt. Deferred because it needs a schema/interface
# change across LLMProvider + both providers + pipeline + scorer +
# repository, and more generated tokens per combo.
#
# If we plan to LLM-review every p2 pass then maybe p2 and p4 should be combined.
#
@@ -80,17 +116,45 @@ You are reviewing a transportation concept for real-world viability — could th
actually be built and operated safely. Whether it is new, exciting, or original
is NOT the question.
## Domain
This concept is being evaluated for "{domain_name}": {domain_description}
Judge every metric against what THIS domain actually needs, not general
expectations for the platform category. A range far beyond what this domain
requires is a strength or a non-issue, never a weakness -- don't reason about
range, speed, or capacity by comparing to what other vehicles of this type
typically have in general use; compare to what this specific domain calls for.
## Concept
{description}
## Metric Scores
All scores below are normalized to 0-1, where HIGHER IS ALWAYS BETTER for
every metric listed, regardless of what the metric measures (this already
accounts for things like "lower cost is better" — you don't need to invert
anything). A score of 1.0 means excellent, not "pegged" or "maxed out badly."
Each metric below is given as its raw estimated physical value (in the unit
shown) AND a normalized score from 0-1, where HIGHER IS ALWAYS BETTER for
every metric listed regardless of what it measures (this already accounts
for things like "lower cost is better" — you don't need to invert anything).
A score of 1.0 means excellent, not "pegged" or "maxed out badly."
Reason from the RAW value first — it's the actual physics. The normalized
score is a summary, not a fact on its own: a real, correct estimate can
still normalize to a low-looking number simply because the domain's scale
was built for a different, more demanding kind of vehicle (a cyclist's real
~5 W/kg legitimately normalizes to ~0.16 next to a car engine's 2000 W/kg
ceiling — that low score doesn't mean the estimate is bad or the concept is
weak, it means human power is small next to a car engine, which everyone
already knows). If a normalized score looks alarming, check whether the raw
value is actually reasonable for what this component fundamentally is
before treating the score as evidence of a problem.
{scores}
Safety and accessibility (infrastructure/regulatory availability) are NOT
among the scores above — neither reduces to a physics formula the way the
metrics above do, so nothing here estimates them numerically. Reason about
both directly from the concept description: does this combination carry a
specific safety hazard, and is the infrastructure/regulatory environment it
needs realistic? Both feed into the RATING below as qualitative judgment
calls, not as scores of their own.
## What makes something IMPLAUSIBLE
Mark IMPLAUSIBLE if either of these is true:
- It is physically or engineering-wise impossible given the components
@@ -102,15 +166,13 @@ Mark IMPLAUSIBLE if either of these is true:
fatiguing a hull over time is a real structural risk, not just "explosives
are dangerous in general"; a fuel that's fine in the open becoming
concentrated in a sealed tube is a real risk, not just "fuel is
flammable"). A low given safety score is a signal the pipeline already
found something concerning — treat it as evidence, not noise to explain
away.
flammable").
- Or: a real regulatory/infrastructure barrier with no plausible workaround.
None of these make something implausible on their own:
- being unoriginal or something like it already exists
- being expensive, slow, or short-range
- a single mediocre score on one metric that isn't safety-related
- a single mediocre score on one metric
Most concepts that reach this review are ordinary and workable; reserve
IMPLAUSIBLE for a real, specific problem you can name — but don't require
@@ -119,14 +181,26 @@ Reason from the physics and engineering actually described here, not from
whether something like it already exists — novelty or lack of it is not
evidence either way.
## Overall Rating
Separately from the plausibility verdict, give ONE holistic rating —
LOW, MEDIUM, or HIGH — for how good this combination is overall. This is a
single combined judgment, not a separate score per attribute: weigh the
metric scores above together with your own qualitative read on safety and
accessibility into one rating, the way a person sizing up the whole concept
would, not a checklist of independent numbers.
## What to write
In 2-4 sentences, give your reasoning, then check it against the scores
above: if your reasoning conflicts with a score (e.g. you believe this is
hazardous but its safety score is high), name the metric and say so
above: if your reasoning conflicts with a score (e.g. you believe cost is
a serious problem but its cost score is high), name the metric and say so
explicitly — don't silently contradict a given score.
Finish with exactly one line:
Finish with exactly two lines. For the first, pick exactly one:
RATING: LOW
RATING: MEDIUM
RATING: HIGH
Then, for the second, pick exactly one:
VERDICT: PLAUSIBLE
or
VERDICT: IMPLAUSIBLE
"""

View File

@@ -11,8 +11,9 @@ from physcom.llm.prompts import (
PHYSICS_ESTIMATION_PROMPT,
PLAUSIBILITY_REVIEW_PROMPT,
format_metrics_for_prompt,
format_scores_for_prompt,
)
from physcom.models.domain import MetricBound
from physcom.models.domain import Domain, MetricBound
class GeminiLLMProvider(LLMProvider):
@@ -46,12 +47,18 @@ class GeminiLLMProvider(LLMProvider):
return parse_metric_json(response.text, metrics)
def review_plausibility(
self, combination_description: str, scores: dict[str, float]
self,
combination_description: str,
raw_metrics: dict[str, float],
normalized_scores: dict[str, float],
domain: Domain,
) -> tuple[str, bool]:
scores_str = "\n".join(f"- {k}: {v:.3f}" for k, v in scores.items())
scores_str = format_scores_for_prompt(raw_metrics, normalized_scores, domain.metric_bounds)
prompt = PLAUSIBILITY_REVIEW_PROMPT.format(
description=combination_description,
scores=scores_str,
domain_name=domain.name,
domain_description=domain.description,
)
try:
response = self._client.models.generate_content(

View File

@@ -3,7 +3,7 @@
from __future__ import annotations
from physcom.llm.base import LLMProvider
from physcom.models.domain import MetricBound
from physcom.models.domain import Domain, MetricBound
class MockLLMProvider(LLMProvider):
@@ -21,9 +21,13 @@ class MockLLMProvider(LLMProvider):
return result
def review_plausibility(
self, combination_description: str, scores: dict[str, float]
self,
combination_description: str,
raw_metrics: dict[str, float],
normalized_scores: dict[str, float],
domain: Domain,
) -> tuple[str, bool]:
avg = sum(scores.values()) / max(len(scores), 1)
avg = sum(normalized_scores.values()) / max(len(normalized_scores), 1)
if avg > 0.5:
return ("This concept appears plausible and worth further investigation.", True)
return ("This concept has significant feasibility challenges.", False)

View File

@@ -12,8 +12,9 @@ from physcom.llm.prompts import (
PHYSICS_ESTIMATION_PROMPT,
PLAUSIBILITY_REVIEW_PROMPT,
format_metrics_for_prompt,
format_scores_for_prompt,
)
from physcom.models.domain import MetricBound
from physcom.models.domain import Domain, MetricBound
class OllamaLLMProvider(LLMProvider):
@@ -34,12 +35,18 @@ class OllamaLLMProvider(LLMProvider):
return parse_metric_json(text, metrics)
def review_plausibility(
self, combination_description: str, scores: dict[str, float]
self,
combination_description: str,
raw_metrics: dict[str, float],
normalized_scores: dict[str, float],
domain: Domain,
) -> tuple[str, bool]:
scores_str = "\n".join(f"- {k}: {v:.3f}" for k, v in scores.items())
scores_str = format_scores_for_prompt(raw_metrics, normalized_scores, domain.metric_bounds)
prompt = PLAUSIBILITY_REVIEW_PROMPT.format(
description=combination_description,
scores=scores_str,
domain_name=domain.name,
domain_description=domain.description,
)
text = self._generate(prompt, json_mode=False).strip()
return (text, parse_verdict(text))
@@ -54,7 +61,7 @@ class OllamaLLMProvider(LLMProvider):
headers={"Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(req, timeout=120) as resp:
with urllib.request.urlopen(req, timeout=300) as resp:
return json.loads(resp.read())["response"]
except urllib.error.URLError as exc:
raise ConnectionError(

View File

@@ -26,6 +26,32 @@ class DomainConstraint:
allowed_values: list[str] = field(default_factory=list) # e.g. ["ground", "air"]
@dataclass
class FreeVariable:
"""A domain-declared quantity pass 2's estimator searches to maximize
the composite score (see Pipeline._estimate_via_formulas), e.g.
"actuator_mass". floor_formula/ceiling_formula are evaluated per-combo
and may reference dep(...) and any free variable declared at a lower
sort_order (mirrors the outer/inner nesting the built-in transport
physics model already does by hand)."""
name: str
floor_formula: str
ceiling_formula: str
sort_order: int = 0
id: int | None = None
@dataclass
class MetricFormula:
"""A domain-declared formula computing one metric's raw value, evaluated
against dep(...) lookups and the domain's resolved free variables."""
metric_name: str
formula: str
id: int | None = None
@dataclass
class Domain:
"""A context frame that defines what 'good' means (e.g., urban_commuting)."""
@@ -34,4 +60,6 @@ class Domain:
description: str = ""
metric_bounds: list[MetricBound] = field(default_factory=list)
constraints: list[DomainConstraint] = field(default_factory=list)
free_variables: list[FreeVariable] = field(default_factory=list)
metric_formulas: list[MetricFormula] = field(default_factory=list)
id: int | None = None

View File

@@ -24,6 +24,7 @@ GROUND_PLATFORMS: list[Entity] = [
Dependency("physical", "mass", "50", "kg", "range_min"),
Dependency("infrastructure", "road_network", "true", None, "requires"),
Dependency("environment", "medium", "ground", None, "requires"),
Dependency("physical", "target_velocity", "25", "m/s", "provides"),
],
),
Entity(
@@ -41,6 +42,7 @@ GROUND_PLATFORMS: list[Entity] = [
Dependency("physical", "mass", "5", "kg", "range_min"),
Dependency("infrastructure", "road_network", "true", None, "requires"),
Dependency("environment", "medium", "ground", None, "requires"),
Dependency("physical", "target_velocity", "6", "m/s", "provides"),
],
),
Entity(
@@ -58,6 +60,7 @@ GROUND_PLATFORMS: list[Entity] = [
Dependency("physical", "mass", "10000", "kg", "range_min"),
Dependency("infrastructure", "rail_network", "true", None, "requires"),
Dependency("environment", "medium", "ground", None, "requires"),
Dependency("physical", "target_velocity", "30", "m/s", "provides"),
],
),
]
@@ -79,6 +82,7 @@ WATER_PLATFORMS: list[Entity] = [
Dependency("physical", "mass", "100000", "kg", "range_max"),
Dependency("physical", "mass", "30", "kg", "range_min"),
Dependency("environment", "medium", "water", None, "requires"),
Dependency("physical", "target_velocity", "8", "m/s", "provides"),
],
),
Entity(
@@ -91,9 +95,11 @@ WATER_PLATFORMS: list[Entity] = [
Dependency("environment", "gravity", "true", None, "provides"),
Dependency("physical", "footprint", "200", "", "range_max"),
Dependency("physical", "footprint", "20", "", "range_min"),
Dependency("physical", "mass", "20000000", "kg", "range_max"),
Dependency("physical", "mass", "10000", "kg", "range_min"),
Dependency("environment", "medium", "water", None, "requires"),
Dependency("physical", "energy_density", "720000", "J/kg", "range_min"),
Dependency("physical", "target_velocity", "8", "m/s", "provides"),
],
),
]
@@ -118,6 +124,7 @@ AIR_PLATFORMS: list[Entity] = [
Dependency("environment", "medium", "air", None, "requires"),
Dependency("physical", "energy_density", "1440000", "J/kg", "range_min"),
Dependency("physical", "min_effective_accel", "2.0", "m/s²", "range_min"),
Dependency("physical", "target_velocity", "60", "m/s", "provides"),
],
),
Entity(
@@ -135,6 +142,7 @@ AIR_PLATFORMS: list[Entity] = [
Dependency("environment", "medium", "air", None, "requires"),
Dependency("physical", "energy_density", "720000", "J/kg", "range_min"),
Dependency("physical", "min_effective_accel", "10", "m/s²", "range_min"),
Dependency("physical", "target_velocity", "30", "m/s", "provides"),
],
),
Entity(
@@ -191,6 +199,20 @@ MULTI_PLATFORMS: list[Entity] = [
Dependency("physical", "footprint", "5", "", "range_min"),
Dependency("physical", "mass", "10000", "kg", "range_max"),
Dependency("physical", "mass", "1500", "kg", "range_min"),
# No requires here previously -- vacuously satisfied every
# domain's medium DomainConstraint (check_domain_constraints
# only flags a violation when an entity DECLARES a requires
# for the constrained key), including space-only
# interplanetary_travel. The current requires/domain-constraint
# model only supports one value per key -- there's no OR
# mechanism for "ground or water" -- so this picks ground
# (its primary, most-common domain) rather than leaving it
# undeclared. Real tradeoff: it can no longer participate in
# maritime_shipping (water-only) either, losing the water half
# of "amphibious." Closes the vacuous-pass hole; genuine
# multi-medium support would need OR semantics added to
# check_domain_constraints, a separate, bigger change.
Dependency("environment", "medium", "ground", None, "requires"),
],
),
]
@@ -214,6 +236,7 @@ FICTIONAL_PLATFORMS: list[Entity] = [
Dependency("physical", "mass", "5000", "kg", "range_min"),
Dependency("infrastructure", "hyperloop_tube", "true", None, "requires"),
Dependency("environment", "medium", "ground", None, "requires"),
Dependency("physical", "target_velocity", "270", "m/s", "provides"), # near-sonic, per its own description
],
),
]
@@ -302,7 +325,7 @@ BIOLOGICAL_ACTUATORS: list[Entity] = [
Dependency("energy", "energy_form", "biological", None, "requires"),
Dependency("physical", "mass", "0", "kg", "range_min"),
Dependency("force", "thrust_profile", "low_continuous", None, "provides"),
Dependency("force", "power_density", "1.5", "W/kg", "provides"),
Dependency("force", "power_density", "5.5", "W/kg", "provides"),
],
),
Entity(
@@ -720,11 +743,29 @@ URBAN_COMMUTING = Domain(
name="urban_commuting",
description="Daily travel within a city, 1-50km range",
metric_bounds=[
# safety and availability removed from the scored/weighted metric set:
# both are judgment calls (risk assessment, infrastructure prevalence),
# not physics quantities with a formula, and running them through the
# same log-normalize() built for physical quantities produced
# incoherent results (a safety raw value already declared as "0-1"
# getting re-normalized into a different, unexplainable 0-1 number --
# see combo 1540's review, where phi4 could only cite the post-
# normalization number with no way to justify it). Safety is now a
# qualitative consideration folded into pass 4's holistic RATING
# instead. Availability needs real per-infrastructure-type research
# this project hasn't done -- not scored anywhere for now rather than
# pretend a quick formula or an equally uninformed LLM guess settles it.
# Weights renormalized to sum to 1.0 across the remaining metrics.
# speed and cargo_capacity_kg added -- a commute's actual travel
# time and whether the vehicle can carry groceries/passengers/gear
# both matter as much as raw power_density did on their own; speed
# is a genuine build OUTPUT (see _raw_physics_from_masses), not a
# platform-declared constant.
MetricBound("power_density", weight=0.25, norm_min=1, norm_max=2000, unit="W/kg"),
MetricBound("cost_efficiency", weight=0.25, norm_min=1e-5, norm_max=2e-3, unit="$/m", lower_is_better=True),
MetricBound("safety", weight=0.25, norm_min=0.0, norm_max=1.0, unit="0-1"),
MetricBound("availability", weight=0.15, norm_min=0.0, norm_max=1.0, unit="0-1"),
MetricBound("speed", weight=0.25, norm_min=2, norm_max=30, unit="m/s"),
MetricBound("range_fuel", weight=0.10, norm_min=5000, norm_max=500000, unit="m"),
MetricBound("cargo_capacity_kg", weight=0.15, norm_min=1, norm_max=500, unit="kg"),
],
constraints=[DomainConstraint("medium", ["ground", "air"])],
)
@@ -733,11 +774,12 @@ INTERPLANETARY = Domain(
name="interplanetary_travel",
description="Travel between planets within a solar system",
metric_bounds=[
MetricBound("power_density", weight=0.30, norm_min=10, norm_max=10000, unit="W/kg"),
MetricBound("range_fuel", weight=0.30, norm_min=1e9, norm_max=1e13, unit="m"),
MetricBound("safety", weight=0.20, norm_min=0.0, norm_max=1.0, unit="0-1"),
MetricBound("cost_efficiency", weight=0.10, norm_min=1.0, norm_max=1e6, unit="$/m", lower_is_better=True),
MetricBound("range_degradation", weight=0.10, norm_min=8640000, norm_max=3.1536e9, unit="s"),
# safety removed -- see URBAN_COMMUTING comment above. Weights
# renormalized across the remaining metrics.
MetricBound("power_density", weight=0.375, norm_min=10, norm_max=10000, unit="W/kg"),
MetricBound("range_fuel", weight=0.375, norm_min=1e9, norm_max=1e13, unit="m"),
MetricBound("cost_efficiency", weight=0.125, norm_min=1.0, norm_max=1e6, unit="$/m", lower_is_better=True),
MetricBound("range_degradation", weight=0.125, norm_min=8640000, norm_max=3.1536e9, unit="s"),
],
constraints=[DomainConstraint("medium", ["space"])],
)
@@ -746,11 +788,12 @@ MARITIME_SHIPPING = Domain(
name="maritime_shipping",
description="Ocean cargo transport between ports, 100-40000km range",
metric_bounds=[
MetricBound("power_density", weight=0.15, norm_min=1, norm_max=1000, unit="W/kg"),
MetricBound("cargo_capacity", weight=0.25, norm_min=1000, norm_max=2e8, unit="kg"),
MetricBound("cost_efficiency", weight=0.25, norm_min=1e-9, norm_max=1e-6, unit="$/(kg\u00b7m)", lower_is_better=True),
MetricBound("safety", weight=0.20, norm_min=0.0, norm_max=1.0, unit="0-1"),
MetricBound("range_fuel", weight=0.15, norm_min=100000, norm_max=40000000, unit="m"),
# safety removed -- see URBAN_COMMUTING comment above. Weights
# renormalized across the remaining metrics.
MetricBound("power_density", weight=0.1875, norm_min=1, norm_max=1000, unit="W/kg"),
MetricBound("cargo_capacity", weight=0.3125, norm_min=1000, norm_max=2e8, unit="kg"),
MetricBound("cost_efficiency", weight=0.3125, norm_min=1e-9, norm_max=1e-6, unit="$/(kg\u00b7m)", lower_is_better=True),
MetricBound("range_fuel", weight=0.1875, norm_min=100000, norm_max=40000000, unit="m"),
],
constraints=[DomainConstraint("medium", ["water"])],
)
@@ -759,11 +802,12 @@ LAST_MILE_DELIVERY = Domain(
name="last_mile_delivery",
description="Short-range package delivery within neighborhoods, 0.5-15km",
metric_bounds=[
MetricBound("power_density", weight=0.25, norm_min=1, norm_max=500, unit="W/kg"),
MetricBound("cost_efficiency", weight=0.30, norm_min=1e-5, norm_max=5e-3, unit="$/m", lower_is_better=True),
MetricBound("cargo_capacity_kg", weight=0.20, norm_min=1, norm_max=500, unit="kg"),
MetricBound("safety", weight=0.15, norm_min=0.0, norm_max=1.0, unit="0-1"),
MetricBound("environmental_impact", weight=0.10, norm_min=0, norm_max=5e-4, unit="kg/m", lower_is_better=True),
# safety removed -- see URBAN_COMMUTING comment above. Weights
# renormalized across the remaining metrics.
MetricBound("power_density", weight=0.2941, norm_min=1, norm_max=500, unit="W/kg"),
MetricBound("cost_efficiency", weight=0.3529, norm_min=1e-5, norm_max=5e-3, unit="$/m", lower_is_better=True),
MetricBound("cargo_capacity_kg", weight=0.2353, norm_min=1, norm_max=500, unit="kg"),
MetricBound("environmental_impact", weight=0.1177, norm_min=0, norm_max=5e-4, unit="kg/m", lower_is_better=True),
],
constraints=[DomainConstraint("medium", ["ground", "air"])],
)
@@ -831,12 +875,11 @@ def load_transport_seed(repo) -> dict:
counts["domains"] += 1
except sqlite3.IntegrityError:
pass
# Backfill metric units and lower_is_better on existing DBs.
for mb in domain.metric_bounds:
repo.ensure_metric(mb.metric_name, unit=mb.unit)
repo.backfill_metric_unit(domain.name, mb.metric_name, mb.unit)
if mb.lower_is_better:
repo.backfill_lower_is_better(domain.name, mb.metric_name)
# Sync domain_metric_weights to exactly match this domain's current
# metric_bounds on existing DBs -- upserts weight/norm_min/norm_max/
# unit for current metrics and removes any that were dropped (e.g.
# safety/availability no longer scored).
repo.sync_domain_metric_weights(domain)
# Backfill domain constraints
repo.replace_domain_constraints(domain)

View File

@@ -6,7 +6,7 @@ from datetime import datetime, timezone
from physcom.db.repository import Repository
from physcom.models.entity import Entity, Dependency
from physcom.models.domain import Domain, DomainConstraint, MetricBound
from physcom.models.domain import Domain, DomainConstraint, FreeVariable, MetricBound, MetricFormula
from physcom.models.combination import Combination
@@ -70,11 +70,27 @@ def export_snapshot(repo: Repository) -> dict:
"key": dc.key,
"allowed_values": dc.allowed_values,
})
fvs = []
for fv in d.free_variables:
fvs.append({
"name": fv.name,
"sort_order": fv.sort_order,
"floor_formula": fv.floor_formula,
"ceiling_formula": fv.ceiling_formula,
})
mfs = []
for mf in d.metric_formulas:
mfs.append({
"metric_name": mf.metric_name,
"formula": mf.formula,
})
domain_list.append({
"name": d.name,
"description": d.description,
"metric_bounds": mbs,
"constraints": dcs,
"free_variables": fvs,
"metric_formulas": mfs,
})
# Export combinations
@@ -207,11 +223,26 @@ def import_snapshot(repo: Repository, data: dict, *, clear: bool = False) -> dic
)
for dc in d_data.get("constraints", [])
]
fvs = [
FreeVariable(
name=fv["name"],
floor_formula=fv["floor_formula"],
ceiling_formula=fv["ceiling_formula"],
sort_order=fv.get("sort_order", 0),
)
for fv in d_data.get("free_variables", [])
]
mfs = [
MetricFormula(metric_name=mf["metric_name"], formula=mf["formula"])
for mf in d_data.get("metric_formulas", [])
]
domain = Domain(
name=d_data["name"],
description=d_data.get("description", ""),
metric_bounds=mbs,
constraints=dcs,
free_variables=fvs,
metric_formulas=mfs,
)
repo.add_domain(domain)
counts["domains"] += 1

View File

@@ -4,7 +4,7 @@ from __future__ import annotations
from flask import Blueprint, flash, redirect, render_template, request, url_for
from physcom.models.domain import Domain, MetricBound
from physcom.models.domain import Domain, FreeVariable, MetricBound, MetricFormula
from physcom_web.app import get_repo
bp = Blueprint("domains", __name__, url_prefix="/domains")
@@ -117,3 +117,96 @@ def metric_delete(domain_id: int, metric_id: int):
flash("Metric removed.", "success")
domain = repo.get_domain_by_id(domain_id)
return render_template("domains/_metrics_table.html", domain=domain)
# ── Free variable CRUD (HTMX partials) ───────────────────────
@bp.route("/<int:domain_id>/free-vars/add", methods=["POST"])
def free_var_add(domain_id: int):
repo = get_repo()
name = request.form["name"].strip()
floor_formula = request.form.get("floor_formula", "").strip()
ceiling_formula = request.form.get("ceiling_formula", "").strip()
try:
sort_order = int(request.form.get("sort_order", "0"))
except ValueError:
sort_order = 0
if not name or not floor_formula or not ceiling_formula:
flash("Name, floor formula, and ceiling formula are required.", "error")
else:
fv = FreeVariable(
name=name, floor_formula=floor_formula, ceiling_formula=ceiling_formula,
sort_order=sort_order,
)
repo.add_free_variable(domain_id, fv)
flash(f"Free variable '{name}' added.", "success")
domain = repo.get_domain_by_id(domain_id)
return render_template("domains/_free_vars_table.html", domain=domain)
@bp.route("/<int:domain_id>/free-vars/<int:fv_id>/edit", methods=["POST"])
def free_var_edit(domain_id: int, fv_id: int):
repo = get_repo()
try:
sort_order = int(request.form.get("sort_order", "0"))
except ValueError:
sort_order = 0
fv = FreeVariable(
name=request.form["name"].strip(),
floor_formula=request.form.get("floor_formula", "").strip(),
ceiling_formula=request.form.get("ceiling_formula", "").strip(),
sort_order=sort_order,
)
repo.update_free_variable(fv_id, fv)
flash("Free variable updated.", "success")
domain = repo.get_domain_by_id(domain_id)
return render_template("domains/_free_vars_table.html", domain=domain)
@bp.route("/<int:domain_id>/free-vars/<int:fv_id>/delete", methods=["POST"])
def free_var_delete(domain_id: int, fv_id: int):
repo = get_repo()
repo.delete_free_variable(fv_id)
flash("Free variable removed.", "success")
domain = repo.get_domain_by_id(domain_id)
return render_template("domains/_free_vars_table.html", domain=domain)
# ── Metric formula CRUD (HTMX partials) ──────────────────────
@bp.route("/<int:domain_id>/formulas/add", methods=["POST"])
def formula_add(domain_id: int):
repo = get_repo()
metric_name = request.form["metric_name"].strip()
formula = request.form.get("formula", "").strip()
if not metric_name or not formula:
flash("Metric name and formula are required.", "error")
else:
repo.add_metric_formula(domain_id, MetricFormula(metric_name=metric_name, formula=formula))
flash(f"Formula for '{metric_name}' added.", "success")
domain = repo.get_domain_by_id(domain_id)
return render_template("domains/_formulas_table.html", domain=domain)
@bp.route("/<int:domain_id>/formulas/<int:formula_id>/edit", methods=["POST"])
def formula_edit(domain_id: int, formula_id: int):
repo = get_repo()
mf = MetricFormula(
metric_name=request.form["metric_name"].strip(),
formula=request.form.get("formula", "").strip(),
)
repo.update_metric_formula(formula_id, mf)
flash("Formula updated.", "success")
domain = repo.get_domain_by_id(domain_id)
return render_template("domains/_formulas_table.html", domain=domain)
@bp.route("/<int:domain_id>/formulas/<int:formula_id>/delete", methods=["POST"])
def formula_delete(domain_id: int, formula_id: int):
repo = get_repo()
repo.delete_metric_formula(formula_id)
flash("Formula removed.", "success")
domain = repo.get_domain_by_id(domain_id)
return render_template("domains/_formulas_table.html", domain=domain)

View File

@@ -32,6 +32,8 @@ def _run_pipeline_in_background(
from physcom.engine.scorer import Scorer
from physcom.engine.pipeline import Pipeline
conn = None
repo = None
try:
conn = init_db(db_path)
repo = Repository(conn)
@@ -58,6 +60,7 @@ def _run_pipeline_in_background(
run_id=run_id,
)
except Exception as exc:
if repo is not None:
try:
repo.update_pipeline_run(
run_id, status="failed",
@@ -65,7 +68,15 @@ def _run_pipeline_in_background(
)
except Exception:
pass
else:
# Couldn't even open the DB to record the failure (bad
# PHYSCOM_DB path, locked/corrupt file) -- the pipeline_runs
# row will stay "pending" forever with no way to write an
# error_message to it, so at least don't let that swallow the
# real cause silently. Server logs are the only trace left.
print(f"pipeline run {run_id} failed before DB was reachable: {exc!r}")
finally:
if conn is not None:
try:
conn.close()
except Exception:

View File

@@ -4,11 +4,25 @@ from __future__ import annotations
from flask import Blueprint, flash, redirect, render_template, request, url_for
from physcom.engine.constraint_resolver import ConstraintResolver
from physcom.engine.pipeline import Pipeline
from physcom.engine.scorer import Scorer
from physcom_web.app import get_repo
bp = Blueprint("results", __name__, url_prefix="/results")
def _run_evaluate(repo, domain, combo, platform_mass=None, actuator_mass=None, storage_mass=None):
"""Purely exploratory -- never writes to the DB. Returns None if this
combo has no free mass allocation to explore (see
Pipeline.evaluate_allocation's docstring)."""
pipeline = Pipeline(repo, ConstraintResolver(), Scorer(domain))
return pipeline.evaluate_allocation(
combo, domain,
platform_mass=platform_mass, actuator_mass=actuator_mass, storage_mass=storage_mass,
)
@bp.route("/")
def results_index():
repo = get_repo()
@@ -25,9 +39,11 @@ def results_domain(domain_name: str):
return redirect(url_for("results.results_index"))
status_filter = request.args.get("status")
results = repo.get_all_results(domain_name, status=status_filter)
rating_filter = request.args.get("rating")
results = repo.get_all_results(domain_name, status=status_filter, rating=rating_filter)
# Domain-scoped status counts (only combos that have results in this domain)
statuses = repo.count_combinations_by_status(domain_name=domain_name)
ratings = repo.count_results_by_rating(domain_name)
return render_template(
"results/list.html",
@@ -35,7 +51,9 @@ def results_domain(domain_name: str):
domain=domain,
results=results,
status_filter=status_filter,
rating_filter=rating_filter,
statuses=statuses,
ratings=ratings,
total_results=sum(statuses.values()),
)
@@ -58,6 +76,7 @@ def result_detail(domain_name: str, combo_id: int):
flash("No results for this combination in this domain.", "error")
return redirect(url_for("results.results_domain", domain_name=domain_name))
scores = repo.get_combination_scores(combo_id, domain.id)
explore_result = _run_evaluate(repo, domain, combo)
return render_template(
"results/detail.html",
@@ -65,6 +84,40 @@ def result_detail(domain_name: str, combo_id: int):
combo=combo,
result=result,
scores=scores,
explore_result=explore_result,
)
@bp.route("/<domain_name>/<int:combo_id>/explore", methods=["POST"])
def explore(domain_name: str, combo_id: int):
"""Live, purely exploratory re-evaluation for an explicit platform/
actuator/storage mass choice -- never touches stored data. Returns an
HTMX partial."""
repo = get_repo()
domain = repo.get_domain(domain_name)
combo = repo.get_combination(combo_id) if domain else None
if not domain or not combo:
return "", 404
def _mass(field: str) -> float | None:
raw = request.form.get(field)
if raw is None:
return None
try:
return float(raw)
except ValueError:
return None
explore_result = _run_evaluate(
repo, domain, combo,
platform_mass=_mass("platform_mass"),
actuator_mass=_mass("actuator_mass"),
storage_mass=_mass("storage_mass"),
)
return render_template(
"results/_explore_result.html",
domain=domain,
explore_result=explore_result,
)
@@ -101,6 +154,7 @@ def submit_review(domain_name: str, combo_id: int):
novelty_flag=novelty_flag,
llm_review=existing.get("llm_review") if existing else None,
human_notes=human_notes,
qualitative_rating=existing.get("qualitative_rating") if existing else None,
)
repo.update_combination_status(combo_id, "reviewed")

View File

@@ -214,6 +214,9 @@ table.compact th, table.compact td { padding: 0.25rem 0.4rem; font-size: 0.83rem
.badge-llm_reviewed { background: rgba(107,163,160,0.12); color: var(--accent-teal); border-color: rgba(107,163,160,0.25); }
.badge-reviewed { background: rgba(155,142,196,0.12); color: var(--accent-violet); border-color: rgba(155,142,196,0.25); }
.badge-pending { background: rgba(184,147,92,0.12); color: var(--accent-amber); border-color: rgba(184,147,92,0.25); }
.badge-rating-low { background: rgba(184,92,92,0.12); color: var(--accent-red); border-color: rgba(184,92,92,0.25); }
.badge-rating-medium { background: rgba(184,147,92,0.12); color: var(--accent-amber); border-color: rgba(184,147,92,0.25); }
.badge-rating-high { background: rgba(122,171,138,0.12); color: var(--accent-green); border-color: rgba(122,171,138,0.25); }
/* ── Buttons ─────────────────────────────────────────────── */
.btn {
@@ -461,6 +464,52 @@ dd { font-size: 0.9rem; color: var(--text-primary); }
margin-left: 0.3rem;
}
/* ── Mass allocation bar (optimizer) ───────────────────────── */
.mass-bar-container {
display: flex;
width: 100%;
height: 18px;
border-radius: 4px;
overflow: hidden;
border: 1px solid var(--border-subtle);
margin-top: 0.5rem;
}
.mass-bar-seg { height: 100%; }
.mass-bar-platform { background: var(--accent-blue); }
.mass-bar-actuator { background: var(--accent-gold); }
.mass-bar-storage { background: var(--accent-teal); }
.mass-bar-legend {
display: flex;
flex-wrap: wrap;
gap: 0.25rem 1rem;
font-size: 0.8rem;
color: var(--text-muted);
margin-top: 0.4rem;
align-items: center;
}
.mass-swatch {
display: inline-block;
width: 10px;
height: 10px;
border-radius: 2px;
margin-right: 0.35rem;
vertical-align: middle;
}
.optimize-summary { margin-bottom: 0.25rem; }
.optimize-score { display: flex; flex-direction: column; gap: 0.1rem; }
/* ── Importance sliders (optimizer) ────────────────────────── */
.weight-slider-row {
display: grid;
grid-template-columns: 140px 1fr 48px;
align-items: center;
gap: 0.75rem;
margin-bottom: 0.5rem;
}
.weight-slider-row label { font-size: 0.85rem; color: var(--text-muted); }
.weight-slider-row output { font-size: 0.85rem; text-align: right; font-variant-numeric: tabular-nums; }
.weight-slider-row input[type="range"] { width: 100%; }
/* ── Select dropdown dark styling ────────────────────────── */
select option {
background: var(--bg-surface);

View File

@@ -0,0 +1,56 @@
<table id="formulas-table">
<thead>
<tr>
<th>Metric</th>
<th>Formula</th>
<th></th>
</tr>
</thead>
<tbody>
{% for mf in domain.metric_formulas %}
<tr>
<td>{{ mf.metric_name }}</td>
<td><code>{{ mf.formula }}</code></td>
<td class="actions">
<button class="btn btn-sm"
onclick="this.closest('tr').nextElementSibling.style.display='table-row'; this.closest('tr').style.display='none'">
Edit
</button>
<form method="post"
hx-post="{{ url_for('domains.formula_delete', domain_id=domain.id, formula_id=mf.id) }}"
hx-target="#formulas-section" hx-swap="innerHTML"
class="inline-form">
<button type="submit" class="btn btn-sm btn-danger">Del</button>
</form>
</td>
</tr>
<tr class="edit-row" style="display:none">
<form method="post"
hx-post="{{ url_for('domains.formula_edit', domain_id=domain.id, formula_id=mf.id) }}"
hx-target="#formulas-section" hx-swap="innerHTML">
<td><input name="metric_name" value="{{ mf.metric_name }}" required></td>
<td><input name="formula" value="{{ mf.formula }}" required></td>
<td>
<button type="submit" class="btn btn-sm btn-primary">Save</button>
<button type="button" class="btn btn-sm"
onclick="this.closest('tr').style.display='none'; this.closest('tr').previousElementSibling.style.display=''">
Cancel
</button>
</td>
</form>
</tr>
{% endfor %}
</tbody>
</table>
<h3>Add Formula</h3>
<form method="post"
hx-post="{{ url_for('domains.formula_add', domain_id=domain.id) }}"
hx-target="#formulas-section" hx-swap="innerHTML"
class="dep-add-form">
<div class="form-row">
<input name="metric_name" placeholder="metric name" required>
<input name="formula" placeholder='formula, e.g. draw_weight_chosen * 2' required>
<button type="submit" class="btn btn-primary">Add</button>
</div>
</form>

View File

@@ -0,0 +1,64 @@
<table id="free-vars-table">
<thead>
<tr>
<th>Name</th>
<th>Order</th>
<th>Floor formula</th>
<th>Ceiling formula</th>
<th></th>
</tr>
</thead>
<tbody>
{% for fv in domain.free_variables %}
<tr>
<td>{{ fv.name }}</td>
<td>{{ fv.sort_order }}</td>
<td><code>{{ fv.floor_formula }}</code></td>
<td><code>{{ fv.ceiling_formula }}</code></td>
<td class="actions">
<button class="btn btn-sm"
onclick="this.closest('tr').nextElementSibling.style.display='table-row'; this.closest('tr').style.display='none'">
Edit
</button>
<form method="post"
hx-post="{{ url_for('domains.free_var_delete', domain_id=domain.id, fv_id=fv.id) }}"
hx-target="#free-vars-section" hx-swap="innerHTML"
class="inline-form">
<button type="submit" class="btn btn-sm btn-danger">Del</button>
</form>
</td>
</tr>
<tr class="edit-row" style="display:none">
<form method="post"
hx-post="{{ url_for('domains.free_var_edit', domain_id=domain.id, fv_id=fv.id) }}"
hx-target="#free-vars-section" hx-swap="innerHTML">
<td><input name="name" value="{{ fv.name }}" required></td>
<td><input name="sort_order" type="number" step="1" value="{{ fv.sort_order }}"></td>
<td><input name="floor_formula" value="{{ fv.floor_formula }}" required></td>
<td><input name="ceiling_formula" value="{{ fv.ceiling_formula }}" required></td>
<td>
<button type="submit" class="btn btn-sm btn-primary">Save</button>
<button type="button" class="btn btn-sm"
onclick="this.closest('tr').style.display='none'; this.closest('tr').previousElementSibling.style.display=''">
Cancel
</button>
</td>
</form>
</tr>
{% endfor %}
</tbody>
</table>
<h3>Add Free Variable</h3>
<form method="post"
hx-post="{{ url_for('domains.free_var_add', domain_id=domain.id) }}"
hx-target="#free-vars-section" hx-swap="innerHTML"
class="dep-add-form">
<div class="form-row">
<input name="name" placeholder="name, e.g. draw_weight_chosen" required>
<input name="sort_order" type="number" step="1" placeholder="order" value="0">
<input name="floor_formula" placeholder='floor formula, e.g. dep("draw_weight", "range_min")' required>
<input name="ceiling_formula" placeholder='ceiling formula, e.g. dep("draw_weight", "range_max")' required>
<button type="submit" class="btn btn-primary">Add</button>
</div>
</form>

View File

@@ -33,4 +33,18 @@
<div id="metrics-section">
{% include "domains/_metrics_table.html" %}
</div>
<h2>Free Variables</h2>
<p class="hint">Quantities pass 2's estimator searches to maximize this domain's composite score. Leave empty for domains where nothing needs sizing.</p>
<div id="free-vars-section">
{% include "domains/_free_vars_table.html" %}
</div>
<h2>Metric Formulas</h2>
<p class="hint">How each metric's raw value is computed from declared entity properties (via <code>dep(key, constraint_type="provides")</code>) and any free variable above. If this domain declares any formulas, pass 2 uses them instead of the built-in vehicle physics model.</p>
<div id="formulas-section">
{% include "domains/_formulas_table.html" %}
</div>
{% endblock %}

View File

@@ -50,12 +50,13 @@
<div class="step-body">
<h3>Physics Estimation</h3>
<p>
Surviving combinations get raw metric estimates &mdash; speed, cost,
safety, range &mdash; via heuristic stubs or an LLM provider that
reasons about the physical properties of each pairing.
Surviving combinations get raw metric estimates &mdash; power
density, cost, range &mdash; from a deterministic physics engine
that sizes each combination from its own declared attributes, not
a guess.
</p>
<div class="step-example">
Bicycle + Human Pedalling &rarr; speed: 20 km/h, cost: $0.01/km
Bicycle + Human Muscle &rarr; power density: 4.4 W/kg, range: 500km
</div>
</div>
</div>
@@ -72,8 +73,8 @@
Combinations are ranked within their domain.
</p>
<div class="step-example">
Domain <code>urban_commuting</code> weights: speed 25%, cost 25%,
safety 25%, availability 15%, range 10%
Domain <code>urban_commuting</code> weights: power density 42%,
cost 42%, range 17%
</div>
</div>
</div>
@@ -85,9 +86,11 @@
<div class="step-body">
<h3>LLM Review</h3>
<p>
Top-scoring combinations are sent to a language model for plausibility
and novelty assessment &mdash; catching physically valid but practically
absurd pairings.
Top-scoring combinations are sent to a language model for a
plausibility verdict plus a holistic LOW/MEDIUM/HIGH rating &mdash;
weighing safety and accessibility as qualitative judgment calls
alongside the physics scores, catching physically valid but
practically absurd pairings.
</p>
<div class="step-example">
"Train + Solar Sail: structurally valid constraints, but solar radiation
@@ -163,14 +166,15 @@
<div class="card concept-card">
<h3>Metrics</h3>
<p>
Quantitative axes like speed, cost, safety, and range. Each metric
has a domain-specific weight and normalization range. Some are
inverted &mdash; lower cost is better.
Quantitative physics axes like power density, cost, and range. Each
metric has a domain-specific weight and normalization range. Some
are inverted &mdash; lower cost is better. Safety and accessibility
are judgment calls, not physics quantities &mdash; they're weighed
qualitatively in the LLM review pass instead of scored here.
</p>
<div class="concept-examples">
<span class="badge">speed</span>
<span class="badge">power_density</span>
<span class="badge">cost_efficiency</span>
<span class="badge">safety</span>
<span class="badge">range_fuel</span>
</div>
</div>

View File

@@ -0,0 +1,56 @@
{% if explore_result is none %}
<p class="empty">No free mass allocation to explore for this combination — its
platform has no declared mass ceiling to bound the sliders.</p>
{% else %}
{% set r = explore_result %}
<div class="optimize-summary">
<div class="optimize-score">
<span class="score-cell" style="font-size:1.4rem">{{ "%.4f"|format(r.composite_score) }}</span>
<span class="subtitle">composite score at this build</span>
</div>
</div>
{% if r.exceeds_platform_envelope %}
<p class="badge badge-p1_fail" style="display:inline-block;margin-bottom:0.75rem">
⚠ total mass {{ "%.1f"|format(r.total_mass) }}kg exceeds this platform's declared ceiling
({{ "%.1f"|format(r.platform_max) }}kg) — not a build this platform category could carry
</p>
{% endif %}
{% if r.insufficient_structure %}
<p class="badge badge-p1_fail" style="display:inline-block;margin-bottom:0.75rem">
⚠ platform mass {{ "%.1f"|format(r.platform_mass) }}kg is too little structure to carry
{{ "%.1f"|format(r.actuator_mass + r.storage_mass) }}kg of actuator+storage
</p>
{% endif %}
<div class="mass-bar-container" title="platform {{ '%.1f'|format(r.platform_mass) }}kg / actuator {{ '%.1f'|format(r.actuator_mass) }}kg / storage {{ '%.1f'|format(r.storage_mass) }}kg">
{% set total = r.total_mass %}
{% if total > 0 %}
<div class="mass-bar-seg mass-bar-platform" style="width: {{ (r.platform_mass / total * 100)|round(1) }}%"></div>
<div class="mass-bar-seg mass-bar-actuator" style="width: {{ (r.actuator_mass / total * 100)|round(1) }}%"></div>
<div class="mass-bar-seg mass-bar-storage" style="width: {{ (r.storage_mass / total * 100)|round(1) }}%"></div>
{% endif %}
</div>
<div class="mass-bar-legend">
<span><span class="mass-swatch mass-bar-platform"></span>platform {{ "%.1f"|format(r.platform_mass) }}kg</span>
<span><span class="mass-swatch mass-bar-actuator"></span>actuator {{ "%.1f"|format(r.actuator_mass) }}kg</span>
<span><span class="mass-swatch mass-bar-storage"></span>storage {{ "%.1f"|format(r.storage_mass) }}kg</span>
<span class="subtitle">{{ "%.1f"|format(r.total_mass) }}kg total</span>
</div>
<table class="compact" style="margin-top:0.75rem">
<thead><tr><th>Metric</th><th>Raw Value</th><th>Normalized</th><th>Weight</th></tr></thead>
<tbody>
{% for mb in domain.metric_bounds %}
{% set val = r.raw_metrics.get(mb.metric_name) %}
{% set n = r.normalized_scores.get(mb.metric_name) %}
<tr>
<td>{{ mb.metric_name }}</td>
<td class="score-cell">{{ val|qty(mb.unit) if val is not none else '—' }}</td>
<td class="score-cell">{{ "%.4f"|format(n) if n is not none else '—' }}</td>
<td>{{ "%.0f%%"|format(mb.weight * 100) }}{{ ' ↓' if mb.lower_is_better else '' }}</td>
</tr>
{% endfor %}
</tbody>
</table>
{% endif %}

View File

@@ -27,6 +27,9 @@
{% if result %}
<dt>Composite Score</dt><dd class="score-cell">{{ "%.4f"|format(result.composite_score) }}</dd>
<dt>Pass Reached</dt><dd>{{ result.pass_reached }}</dd>
{% if result.qualitative_rating %}
<dt>Rating</dt><dd><span class="badge badge-rating-{{ result.qualitative_rating|lower }}">{{ result.qualitative_rating }}</span></dd>
{% endif %}
{% if result.novelty_flag %}
<dt>Novelty</dt><dd>{{ result.novelty_flag }}</dd>
{% endif %}
@@ -102,11 +105,16 @@
{%- elif s.raw_value >= mb.norm_max -%}
<span class="badge badge-{{ 'p1_fail' if mb.lower_is_better else 'valid' }}">at/above max{{ ' (worst)' if mb.lower_is_better else '' }}</span>
{%- else -%}
{% set pct = ((s.raw_value - mb.norm_min) / (mb.norm_max - mb.norm_min) * 100) | int %}
{% set raw_pct = (s.raw_value - mb.norm_min) / (mb.norm_max - mb.norm_min) * 100 %}
{# For lower_is_better metrics, raw_pct alone measures distance from norm_min,
not quality -- a value near norm_min (excellent, cost near its floor) would
otherwise render as a near-empty bar. Invert so the bar and percentage always
mean "how good", matching the normalized score's own higher-is-better convention. #}
{% set pct = ((100 - raw_pct) if mb.lower_is_better else raw_pct) | int %}
<div class="metric-bar-container">
<div class="metric-bar" style="width: {{ pct }}%"></div>
</div>
<span class="metric-bar-label">~{{ pct }}%{{ ' ' if mb.lower_is_better else '' }}</span>
<span class="metric-bar-label">~{{ pct }}%{{ ' (lower is better)' if mb.lower_is_better else '' }}</span>
{%- endif -%}
{%- else -%}
@@ -121,6 +129,53 @@
</div>
{% endif %}
{% if scores %}
<h2>Explore: Scale the Build</h2>
<p class="subtitle">
Purely exploratory — nothing here is saved. Drag a slider to pick a
platform weight class, motor size, or battery size directly, and see how
power density, range, and the resulting score respond. Sliders open on
the saved build above, which is already the score-optimized allocation
for this domain (subject to the platform's physical performance floor),
so the starting point is the best build already found, not an arbitrary
or merely functional one.
</p>
<div class="card">
{% if explore_result is not none %}
{% set r = explore_result %}
<form id="explore-form"
hx-post="{{ url_for('results.explore', domain_name=domain.name, combo_id=combo.id) }}"
hx-trigger="input changed delay:200ms"
hx-target="#explore-result" hx-swap="innerHTML">
<div class="weight-slider-row">
<label for="platform_mass">platform (weight class)</label>
<input type="range" min="{{ r.platform_min }}" max="{{ r.platform_max }}" step="0.1"
id="platform_mass" name="platform_mass" value="{{ r.platform_mass }}"
oninput="document.getElementById('out_platform_mass').textContent = (+this.value).toFixed(1) + 'kg'">
<output id="out_platform_mass">{{ "%.1f"|format(r.platform_mass) }}kg</output>
</div>
<div class="weight-slider-row">
<label for="actuator_mass">actuator (motor/collector size, or operator count/size for muscle power)</label>
<input type="range" min="{{ r.actuator_min }}" max="{{ r.actuator_slider_max }}" step="0.1"
id="actuator_mass" name="actuator_mass" value="{{ r.actuator_mass }}"
oninput="document.getElementById('out_actuator_mass').textContent = (+this.value).toFixed(1) + 'kg'">
<output id="out_actuator_mass">{{ "%.1f"|format(r.actuator_mass) }}kg</output>
</div>
<div class="weight-slider-row">
<label for="storage_mass">storage (battery/tank size)</label>
<input type="range" min="{{ r.storage_min }}" max="{{ r.storage_slider_max }}" step="0.1"
id="storage_mass" name="storage_mass" value="{{ r.storage_mass }}"
oninput="document.getElementById('out_storage_mass').textContent = (+this.value).toFixed(1) + 'kg'">
<output id="out_storage_mass">{{ "%.1f"|format(r.storage_mass) }}kg</output>
</div>
</form>
{% endif %}
<div id="explore-result">
{% include "results/_explore_result.html" %}
</div>
</div>
{% endif %}
<h2>Human Review</h2>
<div id="review-section">
{% include "results/_review_form.html" %}

View File

@@ -26,11 +26,11 @@
{% if statuses %}
<div class="filter-row">
<span>Filter:</span>
<a href="{{ url_for('results.results_domain', domain_name=domain.name) }}"
<span>Status:</span>
<a href="{{ url_for('results.results_domain', domain_name=domain.name, rating=rating_filter) }}"
class="btn btn-sm {{ '' if status_filter else 'btn-primary' }}">All ({{ total_results }})</a>
{% for s, cnt in statuses.items() %}
<a href="{{ url_for('results.results_domain', domain_name=domain.name, status=s) }}"
<a href="{{ url_for('results.results_domain', domain_name=domain.name, status=s, rating=rating_filter) }}"
class="btn btn-sm {{ 'btn-primary' if status_filter == s else '' }}">
{{ s }} ({{ cnt }})
</a>
@@ -38,9 +38,25 @@
</div>
{% endif %}
{% if ratings %}
<div class="filter-row">
<span>Rating:</span>
<a href="{{ url_for('results.results_domain', domain_name=domain.name, status=status_filter) }}"
class="btn btn-sm {{ '' if not rating_filter else 'btn-primary' }}">All</a>
{% for rt in ['HIGH', 'MEDIUM', 'LOW'] %}
{% if rt in ratings %}
<a href="{{ url_for('results.results_domain', domain_name=domain.name, status=status_filter, rating=rt) }}"
class="btn btn-sm {{ 'btn-primary' if rating_filter == rt else '' }}">
{{ rt }} ({{ ratings[rt] }})
</a>
{% endif %}
{% endfor %}
</div>
{% endif %}
{% if not results %}
{% if status_filter %}
<p class="empty">No results with status "{{ status_filter }}" in this domain.</p>
{% if status_filter or rating_filter %}
<p class="empty">No results matching that filter in this domain.</p>
{% else %}
<p class="empty">No results for this domain yet. <a href="{{ url_for('pipeline.pipeline_form') }}">Run the pipeline</a> first.</p>
{% endif %}
@@ -52,6 +68,7 @@
<th>Score</th>
<th>Entities</th>
<th>Status</th>
<th>Rating</th>
<th>Details</th>
<th></th>
</tr>
@@ -69,6 +86,13 @@
<span class="badge badge-{{ r.combination.status }}">{{ r.combination.status }}</span>
{%- endif -%}
</td>
<td>
{%- if r.qualitative_rating -%}
<span class="badge badge-rating-{{ r.qualitative_rating|lower }}">{{ r.qualitative_rating }}</span>
{%- else -%}
{%- endif -%}
</td>
<td class="block-reason-cell">
{%- if r.domain_block_reason -%}
{{ r.domain_block_reason }}

108
tests/test_formula.py Normal file
View File

@@ -0,0 +1,108 @@
"""Tests for the safe formula evaluator."""
import math
import pytest
from physcom.engine.formula import FormulaError, compile_formula, evaluate_formula
class TestArithmetic:
def test_constant(self):
assert evaluate_formula(compile_formula("42"), {}) == 42.0
def test_basic_ops(self):
assert evaluate_formula(compile_formula("2 + 3 * 4"), {}) == 14.0
assert evaluate_formula(compile_formula("(2 + 3) * 4"), {}) == 20.0
assert evaluate_formula(compile_formula("10 / 4"), {}) == 2.5
assert evaluate_formula(compile_formula("2 ** 3"), {}) == 8.0
def test_unary_minus(self):
assert evaluate_formula(compile_formula("-5 + 2"), {}) == -3.0
def test_variable_lookup(self):
result = evaluate_formula(compile_formula("mass * 2"), {"mass": 3.0})
assert result == 6.0
def test_unknown_variable_raises(self):
with pytest.raises(FormulaError):
evaluate_formula(compile_formula("unknown_var"), {})
def test_division_by_zero_raises_formula_error(self):
with pytest.raises(FormulaError):
evaluate_formula(compile_formula("1 / 0"), {})
class TestFunctions:
def test_default_math_functions(self):
assert evaluate_formula(compile_formula("sqrt(16)"), {}) == 4.0
assert evaluate_formula(compile_formula("max(1, 2, 3)"), {}) == 3.0
assert evaluate_formula(compile_formula("min(1, 2, 3)"), {}) == 1.0
assert evaluate_formula(compile_formula("abs(-5)"), {}) == 5.0
assert evaluate_formula(compile_formula("exp(0)"), {}) == 1.0
assert math.isclose(evaluate_formula(compile_formula("log(exp(1))"), {}), 1.0)
def test_custom_injected_function(self):
formula = compile_formula('dep("power_density", "provides")')
result = evaluate_formula(
formula, {}, functions={"dep": lambda key, constraint_type: 99.0}
)
assert result == 99.0
def test_unknown_function_raises(self):
with pytest.raises(FormulaError):
evaluate_formula(compile_formula("unknown_fn(1)"), {})
def test_string_constant_passthrough_to_function(self):
formula = compile_formula('dep("mass")')
result = evaluate_formula(formula, {}, functions={"dep": lambda key: len(key)})
assert result == 4.0
class TestSecurity:
@pytest.mark.parametrize("source", [
"__import__('os').system('echo hi')",
"().__class__",
"[1, 2, 3]",
"{1: 2}",
"{1, 2}",
"(x for x in [1])",
"lambda: 1",
"1 if True else 0",
"1 == 1",
"x.__class__",
"x[0]",
"(lambda: 1)()",
"1; 2",
])
def test_disallowed_constructs_rejected(self, source):
with pytest.raises(FormulaError):
compile_formula(source)
def test_unregistered_function_name_never_executes(self):
"""exec/eval/__import__ etc. parse as ordinary Call nodes -- the
actual guarantee is that no function name is callable unless it's
explicitly in DEFAULT_FUNCTIONS or caller-supplied, checked at
evaluate time, not that the bare name is rejected at compile time."""
with pytest.raises(FormulaError):
evaluate_formula(compile_formula("exec('1')"), {})
def test_dunder_name_rejected(self):
with pytest.raises(FormulaError):
compile_formula("__builtins__")
def test_indirect_call_rejected(self):
with pytest.raises(FormulaError):
compile_formula("(a + b)(1)")
def test_invalid_syntax_raises_formula_error(self):
with pytest.raises(FormulaError):
compile_formula("2 +")
def test_boolean_constant_rejected(self):
with pytest.raises(FormulaError):
compile_formula("True")
def test_large_exponent_overflows_cleanly_not_hangs(self):
with pytest.raises(FormulaError):
evaluate_formula(compile_formula("9 ** 9 ** 9 ** 9"), {})

View File

@@ -69,6 +69,12 @@ def test_blocked_combos_not_scored(seeded_repo):
score_threshold=0.0, passes=[1, 2, 3, 5],
)
# Estimated count should be less than total (blocked ones filtered)
# Estimated count should be less than total (blocked ones filtered).
# Not necessarily equal to pass1_valid + pass1_conditional: a combo can
# pass pass 1's entity-declared-floor checks but still turn out
# structurally infeasible once pass 2 solves the domain-specific
# actuator/storage requirement (e.g. an engine too big to fit its own
# platform's declared mass ceiling) -- that's a legitimate per-domain
# block, not a bug (see Pipeline._decide_masses' `feasible` return).
assert result.pass2_estimated < result.total_generated
assert result.pass2_estimated == result.pass1_valid + result.pass1_conditional
assert result.pass2_estimated <= result.pass1_valid + result.pass1_conditional

View File

@@ -335,28 +335,34 @@ def test_p3_fail_below_threshold(seeded_repo):
def test_p4_fail_implausible(seeded_repo):
"""Combos deemed implausible by LLM should get p4_fail status."""
"""Combos deemed implausible by LLM should get p4_fail status.
Pass 2 is estimator-only now (never calls the LLM), so there's no way
to force every combo's raw estimates toward a controlled low/high value
the way MockLLMProvider's default_estimates used to. Force the pass-4
verdict directly instead -- what's under test here is pipeline.py's
wiring of review_plausibility's return value to status/counters, not
MockLLMProvider's avg-based heuristic.
"""
from physcom.llm.providers.mock import MockLLMProvider
class AlwaysImplausibleLLM(MockLLMProvider):
def review_plausibility(self, description, raw_metrics, normalized_scores, domain):
return ("Always implausible for testing.", False)
repo = seeded_repo
domain = repo.get_domain("urban_commuting")
resolver = ConstraintResolver()
scorer = Scorer(domain)
# Low estimates → normalized scores avg <= 0.5 → MockLLMProvider returns (text, False)
# Use threshold=0.0 so no combo gets p3_fail and all reach pass 4
mock_llm = MockLLMProvider(default_estimates={
"power_density": 0.1, "cost_efficiency": 0.1, "safety": 0.1,
"availability": 0.1, "range_fuel": 0.1,
})
pipeline = Pipeline(repo, resolver, scorer, llm=mock_llm)
pipeline = Pipeline(repo, resolver, scorer, llm=AlwaysImplausibleLLM())
result = pipeline.run(
domain, ["platform", "actuator", "energy_storage"],
score_threshold=0.0, passes=[1, 2, 3, 4],
)
# With low normalized scores (avg <= 0.5), reviewed combos should be p4_fail
assert result.pass4_failed > 0
assert result.pass4_reviewed == 0
@@ -367,20 +373,23 @@ def test_p4_fail_implausible(seeded_repo):
def test_p4_pass_plausible(seeded_repo):
"""Combos deemed plausible by LLM should get llm_reviewed status."""
"""Combos deemed plausible by LLM should get llm_reviewed status.
See test_p4_fail_implausible on why the verdict is forced directly
rather than via controlled pass-2 estimates.
"""
from physcom.llm.providers.mock import MockLLMProvider
class AlwaysPlausibleLLM(MockLLMProvider):
def review_plausibility(self, description, raw_metrics, normalized_scores, domain):
return ("Always plausible for testing.", True)
repo = seeded_repo
domain = repo.get_domain("urban_commuting")
resolver = ConstraintResolver()
scorer = Scorer(domain)
# High estimates → avg > 0.5 → MockLLMProvider returns (text, True)
mock_llm = MockLLMProvider(default_estimates={
"power_density": 500.0, "cost_efficiency": 5e-4, "safety": 0.6,
"availability": 0.7, "range_fuel": 200000.0,
})
pipeline = Pipeline(repo, resolver, scorer, llm=mock_llm)
pipeline = Pipeline(repo, resolver, scorer, llm=AlwaysPlausibleLLM())
result = pipeline.run(
domain, ["platform", "actuator", "energy_storage"],

View File

@@ -0,0 +1,113 @@
"""End-to-end test that pass 2 can estimate a non-transport domain entirely
from domain-authored formulas, without touching the platform/actuator/
energy_storage physics model in Pipeline._estimate_physics."""
import pytest
from physcom.engine.constraint_resolver import ConstraintResolver
from physcom.engine.scorer import Scorer
from physcom.engine.pipeline import Pipeline
from physcom.models.domain import Domain, FreeVariable, MetricBound, MetricFormula
from physcom.models.entity import Dependency, Entity
def _build_archery_domain(repo):
repo.add_entity(Entity(
name="Recurve",
dimension="bow",
dependencies=[
Dependency("physical", "draw_weight", "20", None, "range_min"),
Dependency("physical", "draw_weight", "50", None, "range_max"),
],
))
repo.add_entity(Entity(
name="Carbon",
dimension="arrow",
dependencies=[
Dependency("physical", "arrow_mass", "0.02", None, "provides"),
],
))
return repo.add_domain(Domain(
name="archery_test",
metric_bounds=[
MetricBound("drawback_force", weight=0.6, norm_min=0, norm_max=100),
MetricBound("range", weight=0.4, norm_min=0, norm_max=300),
],
free_variables=[
FreeVariable(
name="draw_weight_chosen",
floor_formula='dep("draw_weight", "range_min")',
ceiling_formula='dep("draw_weight", "range_max")',
sort_order=0,
),
],
metric_formulas=[
MetricFormula(metric_name="drawback_force", formula="draw_weight_chosen * 2"),
MetricFormula(
metric_name="range",
formula='draw_weight_chosen * 5 / dep("arrow_mass")',
),
],
))
def test_formula_domain_scores_without_platform_actuator_shape(repo):
domain = _build_archery_domain(repo)
resolver = ConstraintResolver()
scorer = Scorer(domain)
pipeline = Pipeline(repo, resolver, scorer)
result = pipeline.run(
domain, ["bow", "arrow"], score_threshold=0.01, passes=[1, 2, 3, 5],
)
assert result.total_generated == 1
assert result.pass1_failed == 0
assert result.pass2_estimated == 1
assert result.pass3_above_threshold == 1
combos = repo.list_combinations()
assert len(combos) == 1
combo = combos[0]
scores = {
s["metric_name"]: s["raw_value"]
for s in repo.get_combination_scores(combo.id, domain.id)
}
# Both metrics increase monotonically with draw_weight_chosen and nothing
# trades off against it, so the optimizer should push to the declared
# ceiling (50) -- confirms _search_free_variables is actually searching,
# not just evaluating at the floor.
assert scores["drawback_force"] == pytest.approx(100.0, rel=0.02)
def test_formula_domain_zero_free_variables_direct_evaluation(repo):
"""A domain with metric_formulas but no free_variables should evaluate
each formula once directly -- no search loop at all."""
repo.add_entity(Entity(
name="Recurve",
dimension="bow",
dependencies=[Dependency("physical", "draw_weight", "30", None, "provides")],
))
repo.add_entity(Entity(name="Carbon", dimension="arrow"))
domain = repo.add_domain(Domain(
name="archery_direct_test",
metric_bounds=[MetricBound("drawback_force", weight=1.0, norm_min=0, norm_max=100)],
metric_formulas=[
MetricFormula(metric_name="drawback_force", formula='dep("draw_weight") * 2'),
],
))
resolver = ConstraintResolver()
scorer = Scorer(domain)
pipeline = Pipeline(repo, resolver, scorer)
result = pipeline.run(
domain, ["bow", "arrow"], score_threshold=0.01, passes=[1, 2, 3, 5],
)
assert result.pass2_estimated == 1
combos = repo.list_combinations()
scores = {
s["metric_name"]: s["raw_value"]
for s in repo.get_combination_scores(combos[0].id, domain.id)
}
assert scores["drawback_force"] == 60.0

View File

@@ -1,7 +1,7 @@
"""Tests for the database repository."""
from physcom.models.entity import Entity, Dependency
from physcom.models.domain import Domain, MetricBound
from physcom.models.domain import Domain, FreeVariable, MetricBound, MetricFormula
def test_ensure_dimension(repo):
@@ -63,6 +63,62 @@ def test_add_and_get_domain(repo):
assert loaded.metric_bounds[0].metric_name == "speed"
def test_add_domain_with_free_variables_and_formulas(repo):
domain = Domain(
name="archery_test",
metric_bounds=[MetricBound("drawback_force", weight=1.0, norm_min=0, norm_max=500)],
free_variables=[
FreeVariable(
name="draw_weight",
floor_formula='dep("draw_weight", "range_min")',
ceiling_formula='dep("draw_weight", "range_max")',
sort_order=0,
),
],
metric_formulas=[
MetricFormula(metric_name="drawback_force", formula="draw_weight * 1.5"),
],
)
saved = repo.add_domain(domain)
assert saved.id is not None
loaded = repo.get_domain("archery_test")
assert loaded is not None
assert len(loaded.free_variables) == 1
assert loaded.free_variables[0].name == "draw_weight"
assert loaded.free_variables[0].id is not None
assert len(loaded.metric_formulas) == 1
assert loaded.metric_formulas[0].formula == "draw_weight * 1.5"
def test_free_variable_and_formula_crud(repo):
domain = repo.add_domain(Domain(name="crud_test"))
fv = repo.add_free_variable(
domain.id,
FreeVariable(name="x", floor_formula="0", ceiling_formula="100", sort_order=0),
)
mf = repo.add_metric_formula(
domain.id, MetricFormula(metric_name="m", formula="x * 2")
)
repo.update_free_variable(
fv.id, FreeVariable(name="x", floor_formula="1", ceiling_formula="200", sort_order=0)
)
repo.update_metric_formula(mf.id, MetricFormula(metric_name="m", formula="x * 3"))
loaded = repo.get_domain_by_id(domain.id)
assert loaded.free_variables[0].floor_formula == "1"
assert loaded.free_variables[0].ceiling_formula == "200"
assert loaded.metric_formulas[0].formula == "x * 3"
repo.delete_free_variable(fv.id)
repo.delete_metric_formula(mf.id)
loaded = repo.get_domain_by_id(domain.id)
assert loaded.free_variables == []
assert loaded.metric_formulas == []
def test_combination_save_and_dedup(repo):
e1 = repo.add_entity(Entity(name="A", dimension="platform"))
e2 = repo.add_entity(Entity(name="B", dimension="actuator"))

View File

@@ -7,7 +7,7 @@ import pytest
from physcom.db.schema import init_db
from physcom.db.repository import Repository
from physcom.models.entity import Entity, Dependency
from physcom.models.domain import Domain, DomainConstraint, MetricBound
from physcom.models.domain import Domain, DomainConstraint, FreeVariable, MetricBound, MetricFormula
from physcom.models.combination import Combination
from physcom.snapshot import export_snapshot, import_snapshot
@@ -169,6 +169,46 @@ def test_import_with_combinations(seeded_repo, tmp_path):
assert len(fresh_combos) == len(data["combinations"])
def test_export_import_roundtrip_free_variables_and_formulas(repo, tmp_path):
domain = Domain(
name="archery_snapshot_test",
metric_bounds=[MetricBound("drawback_force", weight=1.0, norm_min=0, norm_max=100)],
free_variables=[
FreeVariable(
name="draw_weight_chosen",
floor_formula='dep("draw_weight", "range_min")',
ceiling_formula='dep("draw_weight", "range_max")',
sort_order=0,
),
],
metric_formulas=[
MetricFormula(metric_name="drawback_force", formula="draw_weight_chosen * 2"),
],
)
repo.add_domain(domain)
data = export_snapshot(repo)
exported = next(d for d in data["domains"] if d["name"] == "archery_snapshot_test")
assert exported["free_variables"] == [{
"name": "draw_weight_chosen", "sort_order": 0,
"floor_formula": 'dep("draw_weight", "range_min")',
"ceiling_formula": 'dep("draw_weight", "range_max")',
}]
assert exported["metric_formulas"] == [
{"metric_name": "drawback_force", "formula": "draw_weight_chosen * 2"},
]
conn = init_db(tmp_path / "fresh.db")
fresh = Repository(conn)
import_snapshot(fresh, data, clear=True)
loaded = fresh.get_domain("archery_snapshot_test")
assert len(loaded.free_variables) == 1
assert loaded.free_variables[0].floor_formula == 'dep("draw_weight", "range_min")'
assert len(loaded.metric_formulas) == 1
assert loaded.metric_formulas[0].formula == "draw_weight_chosen * 2"
def test_import_merge_skips_existing_domain(repo):
"""Merge import skips domains that already exist."""
domain = Domain(