Files
physicalCombinatorics/CLAUDE.md
Andrew Simonson 3795a7e826 rename estimator, skip unused cargo calc, fix six review bugs, close three seed guardrail holes
Rename: _stub_estimate -> _estimate_physics (it's a real deterministic
physics engine now, not a stub) and estimation_method "stub" ->
"physics_calc" to match the value pass 3 already used, for consistency
between the raw-estimate and scored-metric tables. Also skip the
cargo_capacity/cargo_capacity_kg arithmetic entirely in
_raw_physics_from_masses for domains that score neither and don't need
it as cost_efficiency's $/(kg·m) denominator either -- real but modest
savings on the ~11,000-eval-per-combo optimizer hot path (a separate
log1p-caching attempt was tried and reverted: it measured SLOWER, not
faster -- the extra dict lookup cost more than the two math.log1p calls
it avoided).

Six bugs found by a full-codebase review agent, verified individually:

- pipeline.py: LLM rate-limit retry called review_plausibility() with
  domain.metric_bounds instead of domain, crashing the whole pipeline
  run on any retry (every provider immediately accesses domain.name/
  .metric_bounds on that arg).
- _explore_result.html: mass-bar width divided by total_mass with no
  zero guard; biological/ambient actuators can legitimately have 0 mass
  floors, so an all-zero slider combination 500'd the explore endpoint.
- routes/pipeline.py: if init_db/Repository(conn) raised before
  repo/conn were assigned, the except/finally handlers referencing them
  raised UnboundLocalError, silently swallowed by bare except/pass --
  a bad PHYSCOM_DB path left a run stuck at status=pending forever with
  no diagnostic. conn/repo now init to None and are guarded before use;
  the truly-unreachable-DB case at least logs server-side now.
- repository.py: update_combination_status's downgrade guard protected
  scored/llm_reviewed/*_fail but not a write of "valid" -- pass 1
  re-running for a different domain against an already-reviewed combo
  silently reverted its status back to "valid", erasing the review
  signal. Verified directly: marked a combo reviewed, re-ran pass 1,
  status held.
- pipeline.py: cost_efficiency's operating-cost term fell back to
  ground rolling-resistance physics (effective_k_med or ...["ground"])
  for media with no resistance model (space), instead of skipping the
  term the way range_fuel explicitly does two lines above. Every scored
  interplanetary_travel combo got a cost_efficiency computed from
  ground physics applied to a spacecraft. Now reports amortized/upfront
  cost only for such media -- an honest partial answer.
- pipeline.py: `if min_accel and specific_thrust:` used truthiness
  instead of `is not None` -- dep_value() legitimately returns 0.0 for
  a declared floor of zero (Spaceship declares min_effective_accel=0),
  masking a real requirement as "undeclared."

Three seed-data guardrail holes, matching LOGIC DOCS/002's "missing
floor is a silent hole" pattern:

- constraint_resolver.py: CATEGORY_SEVERITY had no entry for the
  "material" category, so Nuclear Thermal Drive/Nuclear Fuel's
  radiation_shielding requirement defaulted to a non-blocking "warn"
  nothing in the catalog ever satisfies. Added material -> block.
  Consequence, verified: every nuclear combo across all domains now
  correctly fails pass 1, since nothing currently provides shielding --
  the accurate state given the catalog gap, not a regression.
- transport_example.py: Submarine had a mass range_min but no
  range_max, unlike its sibling water platform -- _decide_masses skips
  its entire structural-feasibility search when p_max is None. Added a
  20,000,000kg ceiling (small submersible to large ballistic-missile
  class).
- transport_example.py: Amphibious Vehicle declared no medium requires
  at all, so it vacuously satisfied every domain's medium constraint
  including space-only interplanetary_travel. Added medium=ground
  (the current requires model has no OR semantics for "ground or
  water," so this is a real tradeoff -- it can no longer participate in
  maritime_shipping either, losing the water half of "amphibious").
  Verified: interplanetary_travel's pass-2-estimated count dropped from
  33 to 3, and all 3 remaining are genuinely Spaceship-based; the ~30
  removed were confirmed to be Amphibious Vehicle's vacuous passes.

Logged the GPU-batching-for-the-optimizer discussion (why it doesn't
fit at current scale, what threshold would change that, what it would
actually require) as LOGIC DOCS/003 for future reference.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-16 17:33:22 -05:00

81 lines
6.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## What this project is
PhysCom (Physical Combinatorics) — innovation discovery engine that generates entity combinations across dimensions (e.g. platform × power_source), filters by physical constraints, scores against domain-specific metrics, and ranks results. Includes a CLI, a Flask/HTMX web UI, and a 5-pass pipeline (constraints → estimation → scoring → LLM review → human review).
## Commands
- **Install**: `pip install -e ".[dev,web]"` (editable install with test and web deps)
- **Tests (all)**: `python -m pytest tests/ -q` (48 tests, ~5s). Run after every change.
- **Single test file**: `python -m pytest tests/test_scorer.py -q`
- **Single test**: `python -m pytest tests/test_scorer.py::test_score_combination -q`
- **Web dev server**: `python -m physcom_web`
- **CLI**: `python -m physcom` (or `physcom` if installed)
- **Docker**: `docker compose up web` / `docker compose run cli physcom seed`
- **Seed data**: loaded automatically on first DB init (SQLite, `physcom.db` or `$PHYSCOM_DB`)
## Architecture
```
src/physcom/ # Core library (no web dependency)
models/ # Dataclasses: Entity, Dependency, Combination, Domain, MetricBound
db/schema.py # DDL (all CREATE TABLE statements)
db/repository.py # All DB access — single Repository class, sqlite3 row_factory=Row
engine/combinator.py # Cartesian product of entities across dimensions
engine/constraint_resolver.py # Pass 1: requires/excludes/mutex/range/force checks
engine/scorer.py # Pass 3: log-normalize raw→0-1, weighted geometric mean composite
engine/pipeline.py # Orchestrator: combo-first loop, incremental saves, resume, cancel
llm/base.py # LLMProvider ABC (estimate_physics, review_plausibility)
llm/providers/mock.py # MockLLMProvider for tests
seed/transport_example.py # 9 platforms + 9 power sources, 2 domains
src/physcom_web/ # Flask web UI
app.py # App factory, get_repo(), DB path resolution
routes/pipeline.py # Background thread pipeline execution, HTMX status/cancel endpoints
routes/results.py # Results browse, detail view, human review submission
routes/entities.py # Entity CRUD
routes/domains.py # Domain listing
templates/ # Jinja2, extends base.html, uses HTMX for polling
static/style.css # Single stylesheet
tests/ # pytest, uses seeded_repo fixture from conftest.py
```
## Key patterns
- **Guardrails are emergent, not authored.** Entities never reference each other by name. Each declares only its own intrinsic physical envelope (`requires`/`provides`/`excludes` for categorical fit, `range_min`/`range_max` for numeric floors/ceilings on shared keys like `mass`, `footprint`, `energy_density`). `ConstraintResolver` cross-checks these generically across every pair in a combo — a nuclear reactor is blocked from a bicycle because `mass range_min` collides with `mass range_max`, not because any code says "nuclear + bicycle = blocked." When adding physical realism to an entity, add a property to the entity — never a pairwise special case in `ConstraintResolver`. A category missing a floor it should have is a silent guardrail hole (blocks nothing, looks fine). See `LOGIC DOCS/002-guardrails-via-intrinsic-properties.md`.
- **Repository is the only DB interface.** No raw SQL outside `repository.py`.
- **Pipeline is combo-first**: each combo goes through all requested passes before the next combo starts. Progress is persisted per-combo (crash-safe, resumable).
- **`pipeline_runs` table** tracks run lifecycle: pending → running → completed/failed/cancelled. The web route creates the record, then starts a background thread with its own `sqlite3.Connection`.
- **`combination_results`** has rows for ALL combos including blocked ones (pass_reached=1, composite_score=0.0). Scored combos get pass_reached=3+.
- **Status guard**: `update_combination_status` refuses to downgrade `reviewed``scored`.
- **`save_combination`** loads existing status/block_reason on dedup (important for resume).
- **`ensure_metric`** backfills unit if the row already exists with an empty unit.
- **MetricBound** carries `unit` — flows through seed → ensure_metric → metrics table → get_combination_scores → template display.
- **HTMX polling**: `_run_status.html` partial polls every 2s while run is pending/running; stops polling when terminal.
## Data flow (pipeline passes)
1. **Pass 1 — Constraints**: `ConstraintResolver.resolve()` → blocked/conditional/valid. Blocked combos get a result row and `continue`.
2. **Pass 2 — Estimation**: `_estimate_physics()` (deterministic physics engine; estimator-only, no LLM) → raw metric values. Saved immediately via `save_raw_estimates()` (normalized_score=NULL).
3. **Pass 3 — Scoring**: `Scorer.score_combination()` → log-normalized scores + weighted geometric mean composite. Saves via `save_scores()` + `save_result()`.
4. **Pass 4 — LLM Review**: Only for above-threshold combos with an LLM provider. No real provider yet (only `MockLLMProvider`).
5. **Pass 5 — Human Review**: Manual via web UI results page.
## Testing
- Tests use `seeded_repo` fixture (in-memory SQLite with transport seed data: 9 platforms, 9 power sources, 2 domains). There's also a bare `repo` fixture for tests that seed their own data.
- Individual entity fixtures (walking, bicycle, spaceship, solar_sail, etc.) are defined in `conftest.py`.
## Conventions
- Python 3.11+, `from __future__ import annotations` everywhere.
- Dataclasses for models, no ORM.
- Don't use `cd` in Bash commands — run from the working directory so pre-approved permission patterns match.
- Don't add docstrings/comments/type annotations to code you didn't change.
- `INSERT OR IGNORE` won't update existing rows — if adding a new column/field to seed data, also add an UPDATE for backfill.
- Jinja2 `0.0` is falsy — use `is not none` not `if value` when displaying scores that can legitimately be zero.