close guardrail gaps and fix the scoring pipeline top to bottom

Constraint resolver: aggregate mass/footprint across a combo instead of
pairwise-only checks, treat medium/atmosphere as agreement not supply/demand,
reduce multi-provider checks by best/sum instead of AND-ing every provider,
fail closed on unrecognized mutex values, add a propulsion-viability
(thrust-to-weight) rule. Seed data updated to match (nuclear/solar-sail
footprint floors, water-medium exclusions, explicit ground/gravity providers).

Domain metric units were stored globally per metric name instead of
per-domain, silently corrupting cost_efficiency for every domain but the
first one seeded — fixed with a schema migration.

Stub estimator's cost_efficiency/safety/availability/reliability were a
backwards formula and flat constants; replaced with heuristics grounded in
each entity's thrust_profile/energy_form/infrastructure.

LLM estimate_physics() now receives each metric's unit and expected range
instead of a bare name, fixing wildly miscalibrated estimates traced back to
the prompt's own hardcoded example anchoring the model to the wrong order of
magnitude. Sharpened the safety-estimation and plausibility-review prompts.
Deduped provider parsing logic into llm/parsing.py.

Web pipeline form can now pick an LLM provider per run instead of only via
server env var.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-26 00:13:16 -05:00
parent 63295ab80e
commit 434df718d7
18 changed files with 836 additions and 175 deletions

View File

@@ -3,6 +3,7 @@
from __future__ import annotations
from physcom.llm.base import LLMProvider
from physcom.models.domain import MetricBound
class MockLLMProvider(LLMProvider):
@@ -12,11 +13,11 @@ class MockLLMProvider(LLMProvider):
self._defaults = default_estimates or {}
def estimate_physics(
self, combination_description: str, metrics: list[str]
self, combination_description: str, metrics: list[MetricBound]
) -> dict[str, float]:
result = {}
for metric in metrics:
result[metric] = self._defaults.get(metric, 0.5)
for mb in metrics:
result[mb.metric_name] = self._defaults.get(mb.metric_name, 0.5)
return result
def review_plausibility(