Files
physicalCombinatorics/tests/test_llm_parsing.py
Andrew Simonson 434df718d7 close guardrail gaps and fix the scoring pipeline top to bottom
Constraint resolver: aggregate mass/footprint across a combo instead of
pairwise-only checks, treat medium/atmosphere as agreement not supply/demand,
reduce multi-provider checks by best/sum instead of AND-ing every provider,
fail closed on unrecognized mutex values, add a propulsion-viability
(thrust-to-weight) rule. Seed data updated to match (nuclear/solar-sail
footprint floors, water-medium exclusions, explicit ground/gravity providers).

Domain metric units were stored globally per metric name instead of
per-domain, silently corrupting cost_efficiency for every domain but the
first one seeded — fixed with a schema migration.

Stub estimator's cost_efficiency/safety/availability/reliability were a
backwards formula and flat constants; replaced with heuristics grounded in
each entity's thrust_profile/energy_form/infrastructure.

LLM estimate_physics() now receives each metric's unit and expected range
instead of a bare name, fixing wildly miscalibrated estimates traced back to
the prompt's own hardcoded example anchoring the model to the wrong order of
magnitude. Sharpened the safety-estimation and plausibility-review prompts.
Deduped provider parsing logic into llm/parsing.py.

Web pipeline form can now pick an LLM provider per run instead of only via
server env var.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 00:13:16 -05:00

33 lines
1.0 KiB
Python

"""Tests for shared LLM response-parsing logic."""
from __future__ import annotations
from physcom.llm.parsing import parse_metric_json, parse_verdict
from physcom.models.domain import MetricBound
def _bounds():
return [
MetricBound("power_density", weight=0.5, norm_min=1, norm_max=2000, unit="W/kg"),
MetricBound("safety", weight=0.5, norm_min=0.0, norm_max=1.0, unit="0-1"),
]
def test_parse_metric_json_strips_fences():
text = '```json\n{"power_density": 500.0, "safety": 0.7}\n```'
result = parse_metric_json(text, _bounds())
assert result == {"power_density": 500.0, "safety": 0.7}
def test_parse_metric_json_falls_back_to_range_midpoint_on_invalid():
result = parse_metric_json("not json", _bounds())
assert result == {"power_density": 1000.5, "safety": 0.5}
def test_parse_verdict_plausible():
assert parse_verdict("blah blah\nVERDICT: PLAUSIBLE") is True
def test_parse_verdict_implausible():
assert parse_verdict("blah blah\nVERDICT: IMPLAUSIBLE") is False