close guardrail gaps and fix the scoring pipeline top to bottom
Constraint resolver: aggregate mass/footprint across a combo instead of pairwise-only checks, treat medium/atmosphere as agreement not supply/demand, reduce multi-provider checks by best/sum instead of AND-ing every provider, fail closed on unrecognized mutex values, add a propulsion-viability (thrust-to-weight) rule. Seed data updated to match (nuclear/solar-sail footprint floors, water-medium exclusions, explicit ground/gravity providers). Domain metric units were stored globally per metric name instead of per-domain, silently corrupting cost_efficiency for every domain but the first one seeded — fixed with a schema migration. Stub estimator's cost_efficiency/safety/availability/reliability were a backwards formula and flat constants; replaced with heuristics grounded in each entity's thrust_profile/energy_form/infrastructure. LLM estimate_physics() now receives each metric's unit and expected range instead of a bare name, fixing wildly miscalibrated estimates traced back to the prompt's own hardcoded example anchoring the model to the wrong order of magnitude. Sharpened the safety-estimation and plausibility-review prompts. Deduped provider parsing logic into llm/parsing.py. Web pipeline form can now pick an LLM provider per run instead of only via server env var. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
30
src/physcom/llm/parsing.py
Normal file
30
src/physcom/llm/parsing.py
Normal file
@@ -0,0 +1,30 @@
|
||||
"""Shared response-parsing helpers for LLM providers."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import re
|
||||
|
||||
from physcom.models.domain import MetricBound
|
||||
|
||||
|
||||
def parse_verdict(text: str) -> bool:
|
||||
"""Extract VERDICT: PLAUSIBLE/IMPLAUSIBLE from response; default to True."""
|
||||
m = re.search(r"VERDICT:\s*(PLAUSIBLE|IMPLAUSIBLE)", text, re.IGNORECASE)
|
||||
if m:
|
||||
return m.group(1).upper() == "PLAUSIBLE"
|
||||
return True
|
||||
|
||||
|
||||
def parse_metric_json(text: str, metrics: list[MetricBound]) -> dict[str, float]:
|
||||
"""Strip markdown fences and parse JSON; fall back to each metric's own
|
||||
norm_min/norm_max midpoint on error — a flat constant like 0.5 is
|
||||
guaranteed wrong-magnitude for at least some metrics regardless of unit.
|
||||
"""
|
||||
names = {mb.metric_name for mb in metrics}
|
||||
text = re.sub(r"```(?:json)?\s*", "", text).strip().rstrip("`").strip()
|
||||
try:
|
||||
data = json.loads(text)
|
||||
return {k: float(v) for k, v in data.items() if k in names}
|
||||
except (json.JSONDecodeError, ValueError, TypeError):
|
||||
return {mb.metric_name: (mb.norm_min + mb.norm_max) / 2 for mb in metrics}
|
||||
Reference in New Issue
Block a user