Files
physicalCombinatorics/LOGIC DOCS/003-gpu-batching-for-scaled-optimizer.md
Andrew Simonson 3795a7e826 rename estimator, skip unused cargo calc, fix six review bugs, close three seed guardrail holes
Rename: _stub_estimate -> _estimate_physics (it's a real deterministic
physics engine now, not a stub) and estimation_method "stub" ->
"physics_calc" to match the value pass 3 already used, for consistency
between the raw-estimate and scored-metric tables. Also skip the
cargo_capacity/cargo_capacity_kg arithmetic entirely in
_raw_physics_from_masses for domains that score neither and don't need
it as cost_efficiency's $/(kg·m) denominator either -- real but modest
savings on the ~11,000-eval-per-combo optimizer hot path (a separate
log1p-caching attempt was tried and reverted: it measured SLOWER, not
faster -- the extra dict lookup cost more than the two math.log1p calls
it avoided).

Six bugs found by a full-codebase review agent, verified individually:

- pipeline.py: LLM rate-limit retry called review_plausibility() with
  domain.metric_bounds instead of domain, crashing the whole pipeline
  run on any retry (every provider immediately accesses domain.name/
  .metric_bounds on that arg).
- _explore_result.html: mass-bar width divided by total_mass with no
  zero guard; biological/ambient actuators can legitimately have 0 mass
  floors, so an all-zero slider combination 500'd the explore endpoint.
- routes/pipeline.py: if init_db/Repository(conn) raised before
  repo/conn were assigned, the except/finally handlers referencing them
  raised UnboundLocalError, silently swallowed by bare except/pass --
  a bad PHYSCOM_DB path left a run stuck at status=pending forever with
  no diagnostic. conn/repo now init to None and are guarded before use;
  the truly-unreachable-DB case at least logs server-side now.
- repository.py: update_combination_status's downgrade guard protected
  scored/llm_reviewed/*_fail but not a write of "valid" -- pass 1
  re-running for a different domain against an already-reviewed combo
  silently reverted its status back to "valid", erasing the review
  signal. Verified directly: marked a combo reviewed, re-ran pass 1,
  status held.
- pipeline.py: cost_efficiency's operating-cost term fell back to
  ground rolling-resistance physics (effective_k_med or ...["ground"])
  for media with no resistance model (space), instead of skipping the
  term the way range_fuel explicitly does two lines above. Every scored
  interplanetary_travel combo got a cost_efficiency computed from
  ground physics applied to a spacecraft. Now reports amortized/upfront
  cost only for such media -- an honest partial answer.
- pipeline.py: `if min_accel and specific_thrust:` used truthiness
  instead of `is not None` -- dep_value() legitimately returns 0.0 for
  a declared floor of zero (Spaceship declares min_effective_accel=0),
  masking a real requirement as "undeclared."

Three seed-data guardrail holes, matching LOGIC DOCS/002's "missing
floor is a silent hole" pattern:

- constraint_resolver.py: CATEGORY_SEVERITY had no entry for the
  "material" category, so Nuclear Thermal Drive/Nuclear Fuel's
  radiation_shielding requirement defaulted to a non-blocking "warn"
  nothing in the catalog ever satisfies. Added material -> block.
  Consequence, verified: every nuclear combo across all domains now
  correctly fails pass 1, since nothing currently provides shielding --
  the accurate state given the catalog gap, not a regression.
- transport_example.py: Submarine had a mass range_min but no
  range_max, unlike its sibling water platform -- _decide_masses skips
  its entire structural-feasibility search when p_max is None. Added a
  20,000,000kg ceiling (small submersible to large ballistic-missile
  class).
- transport_example.py: Amphibious Vehicle declared no medium requires
  at all, so it vacuously satisfied every domain's medium constraint
  including space-only interplanetary_travel. Added medium=ground
  (the current requires model has no OR semantics for "ground or
  water," so this is a real tradeoff -- it can no longer participate in
  maritime_shipping either, losing the water half of "amphibious").
  Verified: interplanetary_travel's pass-2-estimated count dropped from
  33 to 3, and all 3 remaining are genuinely Spaceship-based; the ~30
  removed were confirmed to be Amphibious Vehicle's vacuous passes.

Logged the GPU-batching-for-the-optimizer discussion (why it doesn't
fit at current scale, what threshold would change that, what it would
actually require) as LOGIC DOCS/003 for future reference.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-16 17:33:22 -05:00

35 lines
4.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GPU batching for the mass-allocation optimizer — not yet, here's the threshold
## Context
`Pipeline._decide_masses`'s joint platform/actuator/storage optimizer (coarse-to-fine grid search, see `_search_best_allocation`) calls its objective function roughly 11,700 times per combo. Profiling confirmed this dominates pipeline runtime: for 50 combos, 582,920 objective-function calls, each doing scalar arithmetic (power_density, the drag cubic solve, normalize, composite_score) on one `(platform, actuator, storage)` triple. The cost is Python's per-call overhead (bytecode dispatch, refcounting, attribute lookups), not the arithmetic itself — the individual formulas are cheap.
## Why GPU doesn't fit today
A single combo's grid is only ~169 points per round (13×13). GPUs pay off when there's enough independent parallel work to amortize kernel-launch and host↔device transfer overhead (each typically tens of microseconds to low milliseconds); 169 elements doesn't come close, and that overhead would be paid repeatedly — once per grid round, ~6-10 rounds per combo.
The parallelism that actually exists is **across combos**, not within one combo's grid — every combo's optimization is fully independent of every other's. At current scale (~180 combos reach the optimizer per domain after Pass 1 filtering), batching every combo's grid into one array gives ~180×169 ≈ 30K elements per round — borderline, probably a wash against plain CPU numpy.
## The actual threshold
Combo count scales **multiplicatively** with added dimensions or entities per dimension (today: 11 platforms × 15 actuators × 18 storages ≈ 2,970 combos, ~180 of which reach the optimizer). Add a 4th dimension with even 10 options and total combos scale to ~30,000, with optimizer-eligible combos likely growing roughly proportionally to ~1,800/domain — batched grid size ≈ 300K elements/round. A 5th dimension does it again, into the low millions. That's the regime where a GPU's thousands of cores start meaningfully outrunning a CPU's 4-16-wide SIMD lanes.
So: not "more dimensions" directly, but the combo×grid batch size those dimensions produce. Rough rule of thumb from this discussion:
- **Tens of thousands of elements/round** (current scale, or a modest one-dimension addition): plain CPU numpy vectorization is enough, no GPU.
- **Hundreds of thousands to low millions**: GPU batching across combos starts being worth evaluating.
## What GPU batching would actually require
Not just "swap numpy for cupy." It means restructuring `_process_pass2` from combo-first (one combo through the optimizer at a time) to batch-first (a chunk of N combos' grids evaluated together as one array with a "combo" axis, broadcasting each combo's own constants — `k_act`, `k_med`, `e_dens`, drag coefficients, mass bounds — across that axis). That's a real architectural change, not a drop-in acceleration:
- **CLAUDE.md documents the pipeline as deliberately combo-first**: "each combo goes through all requested passes before the next combo starts... Progress is persisted per-combo (crash-safe, resumable)." Batching means checkpointing per-*batch*, not per-combo — a real (if manageable) tradeoff against that resumability guarantee, not something that comes free alongside the speedup.
- Every branch in `_raw_physics_from_masses` / `_solve_achievable_speed_mps` (biological floors, ambient energy forms, degenerate fallbacks, the cubic's edge cases) needs to become `np.where(condition, a, b)` instead of `if/else` — careful, error-prone translation work, not mechanical.
## Decision
Don't build this now — no current need at ~180 combos/domain. If dimension count grows enough to matter, do it in two steps:
1. **CPU numpy vectorization first** (batch one combo's grid into arrays, evaluate with vectorized ops instead of a Python double-loop). This is needed regardless of GPU or not, since it's the same rewrite either way, and profiling suggests it could plausibly give 10-50x on its own by replacing ~11,700 Python calls/combo with a couple dozen numpy batch calls.
2. **Re-profile at the new scale.** Only reach for GPU batching-across-combos if CPU numpy is still the dominant cost after step 1, and only once the batched element count is actually in GPU-favorable territory (see thresholds above) — this is a "measure, then decide" call, not something to build ahead of need.