Commit Graph

15 Commits

Author SHA1 Message Date
3795a7e826 rename estimator, skip unused cargo calc, fix six review bugs, close three seed guardrail holes
Rename: _stub_estimate -> _estimate_physics (it's a real deterministic
physics engine now, not a stub) and estimation_method "stub" ->
"physics_calc" to match the value pass 3 already used, for consistency
between the raw-estimate and scored-metric tables. Also skip the
cargo_capacity/cargo_capacity_kg arithmetic entirely in
_raw_physics_from_masses for domains that score neither and don't need
it as cost_efficiency's $/(kg·m) denominator either -- real but modest
savings on the ~11,000-eval-per-combo optimizer hot path (a separate
log1p-caching attempt was tried and reverted: it measured SLOWER, not
faster -- the extra dict lookup cost more than the two math.log1p calls
it avoided).

Six bugs found by a full-codebase review agent, verified individually:

- pipeline.py: LLM rate-limit retry called review_plausibility() with
  domain.metric_bounds instead of domain, crashing the whole pipeline
  run on any retry (every provider immediately accesses domain.name/
  .metric_bounds on that arg).
- _explore_result.html: mass-bar width divided by total_mass with no
  zero guard; biological/ambient actuators can legitimately have 0 mass
  floors, so an all-zero slider combination 500'd the explore endpoint.
- routes/pipeline.py: if init_db/Repository(conn) raised before
  repo/conn were assigned, the except/finally handlers referencing them
  raised UnboundLocalError, silently swallowed by bare except/pass --
  a bad PHYSCOM_DB path left a run stuck at status=pending forever with
  no diagnostic. conn/repo now init to None and are guarded before use;
  the truly-unreachable-DB case at least logs server-side now.
- repository.py: update_combination_status's downgrade guard protected
  scored/llm_reviewed/*_fail but not a write of "valid" -- pass 1
  re-running for a different domain against an already-reviewed combo
  silently reverted its status back to "valid", erasing the review
  signal. Verified directly: marked a combo reviewed, re-ran pass 1,
  status held.
- pipeline.py: cost_efficiency's operating-cost term fell back to
  ground rolling-resistance physics (effective_k_med or ...["ground"])
  for media with no resistance model (space), instead of skipping the
  term the way range_fuel explicitly does two lines above. Every scored
  interplanetary_travel combo got a cost_efficiency computed from
  ground physics applied to a spacecraft. Now reports amortized/upfront
  cost only for such media -- an honest partial answer.
- pipeline.py: `if min_accel and specific_thrust:` used truthiness
  instead of `is not None` -- dep_value() legitimately returns 0.0 for
  a declared floor of zero (Spaceship declares min_effective_accel=0),
  masking a real requirement as "undeclared."

Three seed-data guardrail holes, matching LOGIC DOCS/002's "missing
floor is a silent hole" pattern:

- constraint_resolver.py: CATEGORY_SEVERITY had no entry for the
  "material" category, so Nuclear Thermal Drive/Nuclear Fuel's
  radiation_shielding requirement defaulted to a non-blocking "warn"
  nothing in the catalog ever satisfies. Added material -> block.
  Consequence, verified: every nuclear combo across all domains now
  correctly fails pass 1, since nothing currently provides shielding --
  the accurate state given the catalog gap, not a regression.
- transport_example.py: Submarine had a mass range_min but no
  range_max, unlike its sibling water platform -- _decide_masses skips
  its entire structural-feasibility search when p_max is None. Added a
  20,000,000kg ceiling (small submersible to large ballistic-missile
  class).
- transport_example.py: Amphibious Vehicle declared no medium requires
  at all, so it vacuously satisfied every domain's medium constraint
  including space-only interplanetary_travel. Added medium=ground
  (the current requires model has no OR semantics for "ground or
  water," so this is a real tradeoff -- it can no longer participate in
  maritime_shipping either, losing the water half of "amphibious").
  Verified: interplanetary_travel's pass-2-estimated count dropped from
  33 to 3, and all 3 remaining are genuinely Spaceship-based; the ~30
  removed were confirmed to be Amphibious Vehicle's vacuous passes.

Logged the GPU-batching-for-the-optimizer discussion (why it doesn't
fit at current scale, what threshold would change that, what it would
actually require) as LOGIC DOCS/003 for future reference.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-16 17:33:22 -05:00
3ed3918964 add speed as a derived metric, fix aerodynamic drag gap it exposed, expand urban_commuting
target_velocity was only ever a platform-declared input used to size the
actuator -- an achieved-speed OUTPUT never existed anywhere, even though
trip time clearly matters for a domain like urban commuting. Added
"speed" as a genuine derived metric: achieved steady-state cruise speed
computed from the build's own power_density and the medium's resistance,
the same way power_density/range_fuel/cost_efficiency are already
outputs of a build rather than inputs to it.

That immediately surfaced a known, previously-deferred gap: the
resistance model was mass-proportional only (rolling resistance), with
no velocity-squared aerodynamic drag term, so inverting power/resistance
for speed had no ceiling at all -- light vehicles were "achieving"
thousands of m/s. Added DRAG_POWER_COEFF_BY_MEDIUM (ground only, a
car-like reference cross-section) and a closed-form cubic solve
(_solve_achievable_speed_mps, via Cardano's formula, no iteration) for
the achieved speed where propulsive power balances resistance + drag.
Reused the same effective (drag-inclusive) resistance for range_fuel and
cost_efficiency's operating-cost term, since they're the same physical
quantity (energy spent per meter) evaluated at the build's actual speed.

This also closes the range-overestimation bug flagged much earlier
against combo #876 (a real e-bike): range dropped from ~1,032km to
~53km, right in the ~50-80km realistic e-bike range that was the
original target. Air and water media are unchanged (air's L/D-based
cruise model doesn't have this problem; water hull drag needs its own
treatment, not a car's frontal area -- left as a known remaining gap).

Also added cargo_capacity_kg to urban_commuting (whether a commute
vehicle can carry groceries/passengers/gear matters as much as the
metrics already scored there) and renormalized weights across the now
five metrics.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 19:59:45 -05:00
76f460499a drop safety/availability from scoring, holistic p4 rating, phase-parallel pipeline
safety and availability don't reduce to physics formulas the way
power_density/range_fuel/cost_efficiency do -- they're judgment calls
(risk assessment, infrastructure prevalence), and running them through the
same log-normalize() built for physical quantities produced incoherent
results: safety's raw value is already a "0-1" score, and normalizing it
again turned 0.6 into an unexplainable 0.678 that even the LLM reviewing
it could only cite, never justify (see combo 1540). Removed both from
domain_metric_weights (safety from 4 domains, availability from
urban_commuting) and renormalized the remaining weights to sum to 1.0.

Pass 4 now produces one holistic RATING (LOW/MEDIUM/HIGH) alongside the
existing VERDICT, with safety and accessibility folded in as qualitative
considerations feeding that single judgment rather than scored
separately -- not a checklist of independent numbers. New
qualitative_rating column, filterable in the results UI. Also added
domain name/description to the review prompt so the LLM judges a metric
like range against what the domain actually needs (urban_commuting:
1-50km) instead of generic real-world expectations for the platform
category -- confirmed live on a combo where phi4 had called a 396km range
"limited" by comparing to typical aircraft rather than a domain that
needs 1-50km.

Pass 2 is estimator-only now -- self.llm is never consulted there,
reserved entirely for pass 4. Restructured Pipeline.run() from combo-first
to phase-parallel: each pass now runs to completion across every combo
before the next pass starts, rather than walking each combo through all
four passes before the next combo. This surfaced a real bug: domain-
blocked combos (status stays "valid" by design, not "_fail") were
slipping past a naive status-based skip guard and getting silently
re-processed by pass 2. Fixed with a shared dead-combo check that catches
both generic failures and domain blocks correctly.

Also fixes a results-page display bug found while reviewing a live combo:
the per-metric "position" bar showed raw distance from norm_min without
inverting for lower_is_better metrics, so an excellent cost score (near
the good end) rendered as a ~0%, near-empty bar -- looked bad next to its
own 0.99 normalized score.

Validated live against phi4 (real Ollama calls, not mocked): full-domain
phase-parallel run (2,970 combos, estimator-only p2, 1.6s) followed by a
real pass-4 run (111 reviewed, 11m, 0 crashes, 0 null ratings). Two tests
that relied on the old LLM-driven pass 2 to force deterministic outcomes
were updated to test pass 4's verdict-wiring directly instead. All 100
tests pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 16:52:26 -05:00
be25a837ff replace stub estimator with a requirement-derived physics engine
pass 2's no-LLM fallback previously used broken/placeholder formulas:
power_density passed an actuator's own intensive W/kg straight through
without scaling by vehicle mass, range_fuel multiplied energy_density by a
flat unitless constant, and cost_efficiency was a categorical guess.

Replaces all three with formulas grounded in each entity's own declared
attributes. Actuator and storage mass are sized to what's actually
necessary -- enough power to sustain a platform's target_velocity against
resistance (or real thrust/accel requirements where already declared, for
aircraft/rocket combos), enough energy to reach the domain's own declared
range ceiling -- solved as a closed-form 2x2 linear system rather than an
invented mass-fraction table. Platform mass uses the geometric mean of its
declared range instead of the bare floor, since a category as broad as
Road Vehicle (50kg-36,000kg) is closer to log-uniformly distributed than
uniformly distributed. cost_efficiency is now real operating cost (energy
price x resistance) plus amortized upfront cost (materials cost by medium
x lifetime distance), replacing the old flat per-energy-form guess.

Also fixes two bugs found while validating the above against real-world
reference values: entities whose power source isn't their own carried mass
(Human Muscle, Solar Sail) degenerated to zero power; and combos blocked
by a domain-specific constraint kept combinations.status stuck at "valid"
forever, which silently miscounted them as passing in two repository
queries (count_combinations_by_status, get_pipeline_summary) even though
the per-domain result row was correctly marked blocked.

Adds target_velocity to platforms that had no declared performance
requirement at all, and raises OllamaLLMProvider's HTTP timeout (120s to
300s) to match observed real review-call latency.

Validated against 4 real-world reference combos (commuter car, bicycle,
delivery drone) and a full 2,970-combination domain run (0 pass-2
failures); all 100 existing tests still pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 13:45:01 -05:00
434df718d7 close guardrail gaps and fix the scoring pipeline top to bottom
Constraint resolver: aggregate mass/footprint across a combo instead of
pairwise-only checks, treat medium/atmosphere as agreement not supply/demand,
reduce multi-provider checks by best/sum instead of AND-ing every provider,
fail closed on unrecognized mutex values, add a propulsion-viability
(thrust-to-weight) rule. Seed data updated to match (nuclear/solar-sail
footprint floors, water-medium exclusions, explicit ground/gravity providers).

Domain metric units were stored globally per metric name instead of
per-domain, silently corrupting cost_efficiency for every domain but the
first one seeded — fixed with a schema migration.

Stub estimator's cost_efficiency/safety/availability/reliability were a
backwards formula and flat constants; replaced with heuristics grounded in
each entity's thrust_profile/energy_form/infrastructure.

LLM estimate_physics() now receives each metric's unit and expected range
instead of a bare name, fixing wildly miscalibrated estimates traced back to
the prompt's own hardcoded example anchoring the model to the wrong order of
magnitude. Sharpened the safety-estimation and plausibility-review prompts.
Deduped provider parsing logic into llm/parsing.py.

Web pipeline form can now pick an LLM provider per run instead of only via
server env var.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 00:13:16 -05:00
63295ab80e pre-review fixes 2026-07-25 16:42:39 -05:00
843baa15ad domain-level constraints 2026-03-04 16:53:58 -06:00
00cc8dd9ef I love how stupid this project is
si units and redefining speed metric as thrust/weight ratio
2026-03-04 16:30:09 -06:00
216879bdd5 split actuators from energy storage 2026-03-04 14:18:52 -06:00
e99a14d087 seeding expansion
also: replace energy output with energy output density
2026-03-04 13:21:20 -06:00
f57ac7d6dc QoL and metric value inverter 2026-03-04 11:10:45 -06:00
8dfe3607b1 Add domain CRUD, energy density constraint, LLM status, reset results, score display fixes
Domain management:
- Add domain list/detail/form templates and full CRUD routes (domains.py)
- Add metric bound add/edit/delete via HTMX partials (_metrics_table.html)

Energy density constraint (Rule 6 in ConstraintResolver):
- Hard-block combos where power source provides <25% of platform's required Wh/kg
- Warn (conditional) when under-density but within 4x
- Solar Sail exempt (no stored energy); Airplane requires 400 Wh/kg, Spaceship 2000 Wh/kg
- Add energy_density_wh_kg provides to all 8 stored-energy power sources in seed data
- 3 new constraint resolver tests

LLM-complete status:
- Pipeline Pass 4 now sets combo status to llm_reviewed after successful LLM review
- update_combination_status guards against downgrading: scored won't overwrite
  llm_reviewed or reviewed; llm_reviewed won't overwrite reviewed
- Add badge-llm_reviewed CSS style (light blue)

Reset results:
- Repository.reset_domain_results() deletes combination_results, combination_scores,
  and pipeline_runs for a domain; pipeline re-evaluates on next run
- POST /results/<domain>/reset route with flash confirmation
- "Reset results" danger button with JS confirm dialog in results list

Fix composite score 0 displaying as --- (Jinja2 falsy 0.0 bug):
- Change `if r.composite_score` to `if r.composite_score is not none`

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-19 11:13:00 -06:00
f1b3c75190 Display metric units in web UI, make seed idempotent, simplify code
- Load m.unit in get_domain() so MetricBound carries units from DB
- Add Unit column to domains list template
- Make load_transport_seed() idempotent with IntegrityError handling
  and metric unit backfill for existing DBs
- Remove unused imports (json, sqlite3, Entity)
- Simplify combinator loop to list comprehension
- Merge duplicate conditional/valid branches in pipeline
- Consolidate duplicated SQL in get_all_results()
- Expand CLAUDE.md with fuller architecture docs and conventions

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 16:18:50 -06:00
Simonson, Andrew
d2028a642b Add async pipeline with progress monitoring, resumability, and result transparency
Pipeline engine rewritten with combo-first loop: each combination is processed
through all requested passes before moving to the next, with incremental DB
saves after every step (crash-safe). Blocked combos now get result rows so they
appear in the results page with constraint violation reasons.

New pipeline_runs table tracks run lifecycle (pending/running/completed/failed/
cancelled). Web route launches pipeline in a background thread with its own DB
connection. HTMX polling partial shows live progress with per-pass breakdown.

Also: status guard prevents reviewed->scored downgrade, save_combination loads
existing status on dedup for correct resume, per-metric scores show domain
bounds + units + position bars, ensure_metric backfills units on existing rows.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 15:30:52 -06:00
Simonson, Andrew
8118a62242 Add Flask web UI, Docker Compose, core engine + tests
- physcom core: CLI, 5-pass pipeline, SQLite repo, 37 tests
- physcom_web: Flask app with HTMX for entity/domain/pipeline/results CRUD
- Docker Compose: web + cli services sharing a named volume for the DB
- Clean up local settings to use wildcard permissions

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 13:59:53 -06:00