Pass 4's plausibility review only ever saw normalized scores, never the
raw physical estimate behind them -- confirmed via live testing this was
exactly what caused a real misfire (gemma2:27b cited a real cyclist's
correct 5 W/kg, log-normalized to "0.159" against a car's power scale, as
grounds for rejecting an ordinary bicycle). review_plausibility now takes
raw_metrics + normalized_scores + metric units, and the prompt explicitly
instructs reasoning from the raw value first. Verified live against phi4:
it now cites the actual raw number and correctly explains why a low
normalized score doesn't mean the estimate or concept is bad.
Repository write methods used in the pipeline's hot path now take an
optional commit=False, and Pipeline defers commits during the fast/
deterministic passes (1, 3, and 2 without an LLM), flushing every 200
combos and on any exit path (finally block covers normal completion,
cancellation, and any other exception). LLM-involving calls (pass 2 with
an LLM, all of pass 4) still commit immediately -- those are slow and
crash-prone and worth protecting per-write; the deterministic passes
aren't, and recomputing them is now measured at under a second for the
full domain rather than worth 8,000+ individual fsync'd commits. Full
2,970-combination domain run: multiple minutes -> 0.91s. Test suite:
~70s -> ~15s.
Also fixes two more issues found while auditing the estimator for a real
run: CARGO_KG_PER_STRUCTURAL_KG was 500 (no real vehicle carries 500x its
own structural mass in cargo -- a magnitude bug, not a modeling choice),
corrected to 2.5. And space/rocket platforms' range_fuel now reports the
domain's ceiling instead of an arbitrary placeholder constant -- vacuum
coast isn't resistance-limited, so "distance before running out of fuel"
isn't a meaningful question for these the way it is for ground/air/water
vehicles; the real constraint is delta-v budget, a different metric this
pass doesn't model.
Validated with a live full-domain run (phi4, real Ollama calls): 115
reviewed, 0 malformed/null reviews, 0 verdict-vs-status mismatches.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
pass 2's no-LLM fallback previously used broken/placeholder formulas:
power_density passed an actuator's own intensive W/kg straight through
without scaling by vehicle mass, range_fuel multiplied energy_density by a
flat unitless constant, and cost_efficiency was a categorical guess.
Replaces all three with formulas grounded in each entity's own declared
attributes. Actuator and storage mass are sized to what's actually
necessary -- enough power to sustain a platform's target_velocity against
resistance (or real thrust/accel requirements where already declared, for
aircraft/rocket combos), enough energy to reach the domain's own declared
range ceiling -- solved as a closed-form 2x2 linear system rather than an
invented mass-fraction table. Platform mass uses the geometric mean of its
declared range instead of the bare floor, since a category as broad as
Road Vehicle (50kg-36,000kg) is closer to log-uniformly distributed than
uniformly distributed. cost_efficiency is now real operating cost (energy
price x resistance) plus amortized upfront cost (materials cost by medium
x lifetime distance), replacing the old flat per-energy-form guess.
Also fixes two bugs found while validating the above against real-world
reference values: entities whose power source isn't their own carried mass
(Human Muscle, Solar Sail) degenerated to zero power; and combos blocked
by a domain-specific constraint kept combinations.status stuck at "valid"
forever, which silently miscounted them as passing in two repository
queries (count_combinations_by_status, get_pipeline_summary) even though
the per-domain result row was correctly marked blocked.
Adds target_velocity to platforms that had no declared performance
requirement at all, and raises OllamaLLMProvider's HTTP timeout (120s to
300s) to match observed real review-call latency.
Validated against 4 real-world reference combos (commuter car, bicycle,
delivery drone) and a full 2,970-combination domain run (0 pass-2
failures); all 100 existing tests still pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Constraint resolver: aggregate mass/footprint across a combo instead of
pairwise-only checks, treat medium/atmosphere as agreement not supply/demand,
reduce multi-provider checks by best/sum instead of AND-ing every provider,
fail closed on unrecognized mutex values, add a propulsion-viability
(thrust-to-weight) rule. Seed data updated to match (nuclear/solar-sail
footprint floors, water-medium exclusions, explicit ground/gravity providers).
Domain metric units were stored globally per metric name instead of
per-domain, silently corrupting cost_efficiency for every domain but the
first one seeded — fixed with a schema migration.
Stub estimator's cost_efficiency/safety/availability/reliability were a
backwards formula and flat constants; replaced with heuristics grounded in
each entity's thrust_profile/energy_form/infrastructure.
LLM estimate_physics() now receives each metric's unit and expected range
instead of a bare name, fixing wildly miscalibrated estimates traced back to
the prompt's own hardcoded example anchoring the model to the wrong order of
magnitude. Sharpened the safety-estimation and plausibility-review prompts.
Deduped provider parsing logic into llm/parsing.py.
Web pipeline form can now pick an LLM provider per run instead of only via
server env var.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Domain management:
- Add domain list/detail/form templates and full CRUD routes (domains.py)
- Add metric bound add/edit/delete via HTMX partials (_metrics_table.html)
Energy density constraint (Rule 6 in ConstraintResolver):
- Hard-block combos where power source provides <25% of platform's required Wh/kg
- Warn (conditional) when under-density but within 4x
- Solar Sail exempt (no stored energy); Airplane requires 400 Wh/kg, Spaceship 2000 Wh/kg
- Add energy_density_wh_kg provides to all 8 stored-energy power sources in seed data
- 3 new constraint resolver tests
LLM-complete status:
- Pipeline Pass 4 now sets combo status to llm_reviewed after successful LLM review
- update_combination_status guards against downgrading: scored won't overwrite
llm_reviewed or reviewed; llm_reviewed won't overwrite reviewed
- Add badge-llm_reviewed CSS style (light blue)
Reset results:
- Repository.reset_domain_results() deletes combination_results, combination_scores,
and pipeline_runs for a domain; pipeline re-evaluates on next run
- POST /results/<domain>/reset route with flash confirmation
- "Reset results" danger button with JS confirm dialog in results list
Fix composite score 0 displaying as --- (Jinja2 falsy 0.0 bug):
- Change `if r.composite_score` to `if r.composite_score is not none`
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add LLMRateLimitError to llm/base.py (provider-agnostic)
- GeminiLLMProvider raises it on 429/RESOURCE_EXHAUSTED responses
- Pipeline catches it, marks the run completed (not failed), and returns
partial results — already-reviewed combos are saved, and re-running
pass 4 resumes from where it left off
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add LLMProvider registry (llm/registry.py) that builds a provider from
env vars (LLM_PROVIDER, GEMINI_API_KEY, GEMINI_MODEL)
- Add GeminiLLMProvider using the google-genai SDK
- Wire build_llm_provider() into CLI and web pipeline route (replacing llm=None)
- Wrap pass 2 and pass 4 LLM calls in per-combo try/except so API errors
skip individual combos rather than aborting the whole run
- Add gemini optional dep to pyproject.toml; Dockerfile installs [web,gemini]
- Document env vars in .env.example and README
- Lower requires-python to >=3.10 to match installed system Python
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- physcom core: CLI, 5-pass pipeline, SQLite repo, 37 tests
- physcom_web: Flask app with HTMX for entity/domain/pipeline/results CRUD
- Docker Compose: web + cli services sharing a named volume for the DB
- Clean up local settings to use wildcard permissions
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>