Rename: _stub_estimate -> _estimate_physics (it's a real deterministic physics engine now, not a stub) and estimation_method "stub" -> "physics_calc" to match the value pass 3 already used, for consistency between the raw-estimate and scored-metric tables. Also skip the cargo_capacity/cargo_capacity_kg arithmetic entirely in _raw_physics_from_masses for domains that score neither and don't need it as cost_efficiency's $/(kg·m) denominator either -- real but modest savings on the ~11,000-eval-per-combo optimizer hot path (a separate log1p-caching attempt was tried and reverted: it measured SLOWER, not faster -- the extra dict lookup cost more than the two math.log1p calls it avoided). Six bugs found by a full-codebase review agent, verified individually: - pipeline.py: LLM rate-limit retry called review_plausibility() with domain.metric_bounds instead of domain, crashing the whole pipeline run on any retry (every provider immediately accesses domain.name/ .metric_bounds on that arg). - _explore_result.html: mass-bar width divided by total_mass with no zero guard; biological/ambient actuators can legitimately have 0 mass floors, so an all-zero slider combination 500'd the explore endpoint. - routes/pipeline.py: if init_db/Repository(conn) raised before repo/conn were assigned, the except/finally handlers referencing them raised UnboundLocalError, silently swallowed by bare except/pass -- a bad PHYSCOM_DB path left a run stuck at status=pending forever with no diagnostic. conn/repo now init to None and are guarded before use; the truly-unreachable-DB case at least logs server-side now. - repository.py: update_combination_status's downgrade guard protected scored/llm_reviewed/*_fail but not a write of "valid" -- pass 1 re-running for a different domain against an already-reviewed combo silently reverted its status back to "valid", erasing the review signal. Verified directly: marked a combo reviewed, re-ran pass 1, status held. - pipeline.py: cost_efficiency's operating-cost term fell back to ground rolling-resistance physics (effective_k_med or ...["ground"]) for media with no resistance model (space), instead of skipping the term the way range_fuel explicitly does two lines above. Every scored interplanetary_travel combo got a cost_efficiency computed from ground physics applied to a spacecraft. Now reports amortized/upfront cost only for such media -- an honest partial answer. - pipeline.py: `if min_accel and specific_thrust:` used truthiness instead of `is not None` -- dep_value() legitimately returns 0.0 for a declared floor of zero (Spaceship declares min_effective_accel=0), masking a real requirement as "undeclared." Three seed-data guardrail holes, matching LOGIC DOCS/002's "missing floor is a silent hole" pattern: - constraint_resolver.py: CATEGORY_SEVERITY had no entry for the "material" category, so Nuclear Thermal Drive/Nuclear Fuel's radiation_shielding requirement defaulted to a non-blocking "warn" nothing in the catalog ever satisfies. Added material -> block. Consequence, verified: every nuclear combo across all domains now correctly fails pass 1, since nothing currently provides shielding -- the accurate state given the catalog gap, not a regression. - transport_example.py: Submarine had a mass range_min but no range_max, unlike its sibling water platform -- _decide_masses skips its entire structural-feasibility search when p_max is None. Added a 20,000,000kg ceiling (small submersible to large ballistic-missile class). - transport_example.py: Amphibious Vehicle declared no medium requires at all, so it vacuously satisfied every domain's medium constraint including space-only interplanetary_travel. Added medium=ground (the current requires model has no OR semantics for "ground or water," so this is a real tradeoff -- it can no longer participate in maritime_shipping either, losing the water half of "amphibious"). Verified: interplanetary_travel's pass-2-estimated count dropped from 33 to 3, and all 3 remaining are genuinely Spaceship-based; the ~30 removed were confirmed to be Amphibious Vehicle's vacuous passes. Logged the GPU-batching-for-the-optimizer discussion (why it doesn't fit at current scale, what threshold would change that, what it would actually require) as LOGIC DOCS/003 for future reference. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
4.3 KiB
GPU batching for the mass-allocation optimizer — not yet, here's the threshold
Context
Pipeline._decide_masses's joint platform/actuator/storage optimizer (coarse-to-fine grid search, see _search_best_allocation) calls its objective function roughly 11,700 times per combo. Profiling confirmed this dominates pipeline runtime: for 50 combos, 582,920 objective-function calls, each doing scalar arithmetic (power_density, the drag cubic solve, normalize, composite_score) on one (platform, actuator, storage) triple. The cost is Python's per-call overhead (bytecode dispatch, refcounting, attribute lookups), not the arithmetic itself — the individual formulas are cheap.
Why GPU doesn't fit today
A single combo's grid is only ~169 points per round (13×13). GPUs pay off when there's enough independent parallel work to amortize kernel-launch and host↔device transfer overhead (each typically tens of microseconds to low milliseconds); 169 elements doesn't come close, and that overhead would be paid repeatedly — once per grid round, ~6-10 rounds per combo.
The parallelism that actually exists is across combos, not within one combo's grid — every combo's optimization is fully independent of every other's. At current scale (~180 combos reach the optimizer per domain after Pass 1 filtering), batching every combo's grid into one array gives ~180×169 ≈ 30K elements per round — borderline, probably a wash against plain CPU numpy.
The actual threshold
Combo count scales multiplicatively with added dimensions or entities per dimension (today: 11 platforms × 15 actuators × 18 storages ≈ 2,970 combos, ~180 of which reach the optimizer). Add a 4th dimension with even 10 options and total combos scale to ~30,000, with optimizer-eligible combos likely growing roughly proportionally to ~1,800/domain — batched grid size ≈ 300K elements/round. A 5th dimension does it again, into the low millions. That's the regime where a GPU's thousands of cores start meaningfully outrunning a CPU's 4-16-wide SIMD lanes.
So: not "more dimensions" directly, but the combo×grid batch size those dimensions produce. Rough rule of thumb from this discussion:
- Tens of thousands of elements/round (current scale, or a modest one-dimension addition): plain CPU numpy vectorization is enough, no GPU.
- Hundreds of thousands to low millions: GPU batching across combos starts being worth evaluating.
What GPU batching would actually require
Not just "swap numpy for cupy." It means restructuring _process_pass2 from combo-first (one combo through the optimizer at a time) to batch-first (a chunk of N combos' grids evaluated together as one array with a "combo" axis, broadcasting each combo's own constants — k_act, k_med, e_dens, drag coefficients, mass bounds — across that axis). That's a real architectural change, not a drop-in acceleration:
- CLAUDE.md documents the pipeline as deliberately combo-first: "each combo goes through all requested passes before the next combo starts... Progress is persisted per-combo (crash-safe, resumable)." Batching means checkpointing per-batch, not per-combo — a real (if manageable) tradeoff against that resumability guarantee, not something that comes free alongside the speedup.
- Every branch in
_raw_physics_from_masses/_solve_achievable_speed_mps(biological floors, ambient energy forms, degenerate fallbacks, the cubic's edge cases) needs to becomenp.where(condition, a, b)instead ofif/else— careful, error-prone translation work, not mechanical.
Decision
Don't build this now — no current need at ~180 combos/domain. If dimension count grows enough to matter, do it in two steps:
- CPU numpy vectorization first (batch one combo's grid into arrays, evaluate with vectorized ops instead of a Python double-loop). This is needed regardless of GPU or not, since it's the same rewrite either way, and profiling suggests it could plausibly give 10-50x on its own by replacing ~11,700 Python calls/combo with a couple dozen numpy batch calls.
- Re-profile at the new scale. Only reach for GPU batching-across-combos if CPU numpy is still the dominant cost after step 1, and only once the batched element count is actually in GPU-favorable territory (see thresholds above) — this is a "measure, then decide" call, not something to build ahead of need.