Skip to content

Close the fpga_gate and asic_gate: core/vm timing, dxa drain, baseline refresh - #419

Merged
tinebp merged 9 commits into
masterfrom
fix/verilator-prebuilt-parts
Sep 27, 2026
Merged

tinebp merged 9 commits into
masterfrom
fix/verilator-prebuilt-parts

Conversation

@tinebp

@tinebp tinebp commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator

This branch makes both synthesis gates pass on every DUT. Each commit fixes one root cause, and none of them moves a cycle count.

Gate Result Before
fpga_gate (xcu55c) 12/12 PASS core and vm failed their Fmax floor, top failed its LUT cap, tensor was STALE
asic_gate (Yosys/OpenSTA, ASAP7) 8/8 PASS dxa over its area limit and STALE

Commits

Commit Area Fix Measured
fe67ae44b config The LMEM floor now holds a full warp-count CTA group. At 8 KiB, an NT=4 core fit only three of the four co-resident CTAs that sgemm2_dxa_mcast waits on, so SimX and RTL both hung. rtlsim matches the recorded baseline; model_parity 29/29
26f4882d6 dxa The LMEM drain is back to one element per beat. The word-granular gather cost ~6.9k LUT of mux fabric on an unpipelinable path. Xilinx 249 → 305 MHz, LUT +41.6% → +2%; ASAP7 cells +14% → +2%. Costs ~18% on DXA-fed WGMMA shapes
727b8710b ci Re-records the asic_gate dxa baseline, which was the last one left on the old Yosys/OpenSTA toolchain. STALE → PASS
a7888b314 core Takes the issue arbiter out of the scoreboard hazard loop and narrows the stack AGU. The NT16 dispatch queue and LSU request queue get a default-off, FPGA-only BRAM hint. core 272.8 → 285.2 MHz
87beeb8b6 cache, vm L2 AMO engine: registered look-ahead lane hits; coalesce computed for both drain outcomes; settling entry kept disjoint so the forward merge is an OR-tree; response reuses the bank's forwarded word. TLB: per-entry permission check. dmmu output pipe: new opt-in LATE_VALID skid mode in VX_stream_buffer. L2 fill crossbar: OUT_BUF 2. vm 234.2 → 250.4 MHz, LUT −12.4%; top LUT +5.4% → +4.5%; cache LUT −5.8%
d5a47950e core CSR decode and read-modify-write run in the one-cycle wait every CSR request already takes (for the CTA RAM). Only values that can still move are read at fire, and trap-CSR writes use a one-hot target. Also fixes two latent bugs: CSR writes now happen on fire (a full response buffer made CSRRW return the new value), and VX_dcr_data latches the MPM class with its address. core 268.5 → 286 MHz; SFU lane dispatch about −300 LUT/core
428d65c11 fpu At LATENCY 5, fcvt counts leading zeros on ~x for a negative integer, plus a power-of-two correction, instead of waiting on the negate. Dispatcher → fcvt path; identity checked exhaustively over int32
248fb33b6 ci Records the improved cache and vm synthesis baselines. —
feb8965ef ci Re-records tensor. Pinning its impl_strategy (ebc67f7a9, ce9bb0c55) changed its config hash without a refresh, so it had read STALE since 08-28. STALE → PASS; within tolerance of the old numbers

Verification

  • fpga_gate: full 12-DUT run on the final tree, all PASS:

    DUT Fmax (MHz) LUT
    core 286 179,058
    vm 250.4 119,695
    top 253 133,027 (+4.5%)
    cache 302 28,369
    tensor 250.5 239,260
    gfx, tcu, tex, rtu, dxa, raster, om unchanged unchanged
  • asic_gate: full 8-DUT run, all PASS, every DUT within ±1% of its baseline.

  • rtlsim, cycle- and instruction-identical to pre-change runs:

    • core: 34/34 perf lines.
    • amo + vm: 31/31 passing.
    • riscv ISA: 11/11 passing, including isa-32f-dsp, the only config that exercises the RTL fcvt. The default rtlsim FPU is DPI.
  • New SIMULATION assertions: each new registered or look-ahead path is checked against the logic it replaces. Each check was confirmed to fire on a known-bad mutation of its own path:

    • CSR capture
    • AMO registered hits and settling-entry disjointness
    • AMO in-place old value
    • fcvt LZC
  • Libraries: the VX_stream_buffer LATE_VALID mode is off by default. It matched the accept-enabled form over 4M random cycles, and a known-bad variant fails.

  • No SimX change: no cycle count moved. That is covered by the rtlsim identity above, plus model_parity for the config and core commits.

Known, not addressed here

  • Stale perf baselines: 16 core rtlsim cases fail perf_gate only because their perf-baseline config hashes are stale. The failing set is identical before and after this branch, and regenerating them is a separate reviewed step. The three WGMMA+DXA perf baselines (see 26f4882d6) are in the same state.
  • Thin margins: core is 2.8 MHz above its 283.2 MHz floor, and top is 0.5% under its LUT cap.
  • Trap-CSR ordering race: csrw mepc; mret and csrw mtvec; ecall are not ordered, because CSR ops are not wstall. No current test reaches it, including the rv32ui-p startup, which does exactly csrw mepc; mret. It needs a directed test and a matching SimX decode change.
  • asic_gate toolchain path: build32_asic_gate must be run with TOOLDIR=~/tools; its configured TOOLDIR points at a path that no longer exists.

🤖 Generated with Claude Code

tinebp and others added 9 commits September 26, 2026 16:27
The shared-memory floor is the capacity a core keeps when bank count alone
would leave it too small. At 8192 bytes a default NT=4 core holds only three
of the four co-resident CTAs that `sgemm2_dxa_mcast` rendezvouses on, so its
group barrier can never be satisfied and the launch makes no progress: both
SimX and RTL hang, with no diagnostic from either.

The floor now holds a whole warp-count group. Configurations at NT>=8 already
derive more capacity from `LMEM_NUM_BANKS * lmem_bytes_per_bank` and are
unchanged; NT<=4 regains the window it needs.

rtlsim reproduces the recorded baseline for the case at 14344887 cycles
against 14344871, and the full model_parity cell passes 29/29 with the two
tracked known-issue cases behaving as annotated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The word-granular gather assembled a whole LMEM word per beat: a 512-bit
window into the fill buffer, a six-stage spread network placing the run at
the element stride, and a second six-stage barrel shifting it to the in-word
offset. That is roughly 6.9k LUT of mux fabric in a unit that totals 10k, and
twelve mux levels land in one path whose source is the drain-address
recurrence, so it cannot be pipelined without giving back the throughput it
exists to provide.

The cost broke both synthesis gates: the Xilinx DUT fell to 249.3 MHz against
a 300 MHz target (-17.7% on its baseline) at 14300 LUT (+41.6%), and the
ASAP7 DUT grew to 133033 cells (+14%). The ASIC growth is the RTL alone --
rebuilding at the baseline's own toolchain reproduces it to within 1 cell --
so neither figure was a stale-baseline artifact.

Draining one element per beat restores 305.4 MHz at 10260 LUT (+2%) on
Xilinx, and 118381 cells (+2%) on ASAP7. The multicast and warp-group parity
cells regain the known-issue marks that the wider drain had retired.

This gives back a measured ~18% on the DXA-fed WGMMA shapes. Two ways to keep
the gather were tried and measured first: folding the in-word offset into the
read pointer, which cancels out to the same mux depth, and registering the
write at the worker boundary, which bought 10 MHz for 631 LUT -- worse on the
metric that is gated.

The three WGMMA+DXA perf baselines still describe the wider drain and need
regenerating; they are left untouched rather than edited, since the workload
has moved since they were recorded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The dxa entry was the one asic_gate build left on Yosys 0.57+157 / OpenSTA 2.7.0 when the rest moved to 0.69 / 3.1.0, so it read STALE on every nightly and its verdict said nothing about the design.

Recorded at 1068.6 MHz, 12449 um^2, 118381 cells.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ack AGU

The core fpga_gate DUT (NT16/NW16, xcu55c) had fallen from 298.1 MHz to
272.8 MHz. Aggregating the violating paths of every failing build put
183 of 300 in VX_scoreboard, the rest in the LSU request path; two
causes, both latency-neutral to fix.

Scoreboard: operands_ready_n was computed after the issue decision, so
operands_ready_r -> 16-way out_arb grant -> staging_fire -> reserve rd
and pick the next instr's mask -> NUM_REGS-wide reduce -> operands_ready_r
all fit in one cycle. Readiness is now evaluated for both outcomes (staging
holds / staging fires) from grant-independent state -- one reduction per
candidate instr plus a register compare for the reserved rd -- and the
grant only selects. Commit also registers a per-warp one-hot eop_wis ahead
of its output buffer, so a warp's release reads its own flop instead of
every warp decoding one fanout-52 wis.

Stack AGU: the stack interleave put a subtract, a permute and an add,
three dependent full-width CPAs plus two full-width compares, between the
LSU lane dispatch and the request queue. The permute stays inside the group
field, so the base cancels on the high bits; what remains is a narrow
correction the width of the group field above the window base's trailing
zeros. Plain and pack-load forms share one adder, and the window test
compares only bits above the stack size.

Verified:
- fpga_gate core, at the tip of this series: 300.8 MHz, 178,623 LUT
  (+4.0%) -- PASS. top, tensor and gfx unchanged within tolerance.
- core rtlsim: 25/25 runs cycle- and instruction-identical to before,
  9/9 model_parity pass. No SimX change: no cycle moved.
- AGU equivalence vs the direct remap: 168 configs, 6.5M addresses,
  including non-power-of-two core counts; 0 mismatches.
- hw/unittest issue, mem_scheduler, generic_queue, core: clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The vm fpga_gate build missed its 237.5 MHz floor (234 MHz). Its worst
paths all ended at the L2 bank's core response queue, the dmmu output
stage and the L2 fill crossbar. Every change below is cycle-exact.

L2 AMO engine:
- With a deferred commit stage, the write-back queue's lane hits are
  registered from the next commit key, folding in slot occupancy, so the
  forward network and the response start from flops.
- Coalesce and slot selection are computed for both drain outcomes, so
  the late wb_fire only picks between them.
- A queued result drops its bytes from the settling entry. Each byte then
  has at most one writer, and the forward merge is an OR-tree instead of
  a priority mux.
- The in-place AMO response is the bank's forwarded word under the byte
  mask, which takes the lane select off the response path.
- The bank's S0 atomic hit prediction is the raw tag match. Atomics are
  exempt from the hit-order hazard, and this keeps the MSHR probe off
  the commit_busy cone.

dmmu:
- The TLB permission check runs per entry beside the tag compare, rather
  than behind the selected entry's flags.
- The park lane is a one-hot select of the lane categories.

L2 fill crossbar: OUT_BUF 2 registers the payload at the input
handshake, so it no longer loads on the bank's late ready. Latency is
unchanged.

Results:
- fpga_gate, at the tip of this series: vm 250.0 MHz (was 234.2),
  LUT 119,682 (-12.4%). top 250.0 MHz, LUT +4.5% (was +5.4%, over the
  cap); the dedup of the forwarded word in the AMO engine accounts for
  that recovery. cache 301.8 MHz, LUT 28,369 (-5.8%).
- rtlsim amo+vm: 31/31, cycle- and instruction-identical to the
  pre-change run; amo base/threads8 and vm tlb-stress-l1 re-run at the
  series tip.
- The SIMULATION checks on the registered hits and on the settling-entry
  disjointness fire on known-bad mutations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The core fpga_gate build fell below its 283.2 MHz floor (268.5 MHz). The
worst paths ran from SFU lane dispatch through the 12-bit CSR address
decode, the scheduler's per-warp trap-CSR muxes and the lane
replication, into the CSR response buffer and back into the scheduler's
trap registers. Every CSR request is already held one cycle so the CTA
context RAM can be read. That cycle now does the decode.

Read and write path:
- The stable value, a one-hot of the sources that can still move
  (counters, thread/warp masks, fcsr, the CTA RAM), the write target and
  the read-modify-write result are all registered in the wait cycle.
- The fire cycle is a small AND-OR, and the trap-CSR write is a one-hot
  target, so the scheduler compares no addresses.
- Per-lane thread ids are the lane index on top of a registered base.
- A SIMULATION check compares every registered value against its live
  form at fire.

Host counter reads get their own MPM-window decode and no longer wait on
a CSR request. The kernel-path decode drops the MPM tree, which always
read as zero for class 0.

Two latent bugs are fixed along the way:
- CSR writes are qualified by fire, not by the request being valid. With
  the response buffer full, a CSRRW used to write early and then return
  the new value.
- VX_dcr_data latches the MPM class with its address, instead of reading
  it live off the DCR bus after the request has left.

Results:
- fpga_gate: core 286 MHz (floor 283.2). The CSR paths left the worst
  100; SFU lane dispatch shrinks by about 300 LUTs per core.
- rtlsim core: all 34 perf lines are cycle- and instruction-identical to
  the pre-change run. The 16 failures there are the pre-existing stale
  perf-baseline config hashes, the same set before and after.
- The capture assertion fires on a known-bad mutation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After the CSR fix, core's worst path ran from the dispatcher into
fcvt. At LATENCY 5 (XLEN=32), the integer negate and the LZC share a
cycle, fed across a long route from the dispatcher's output register.

The LZC now reads the one's complement y = ~x of a negative integer.
|x| = y + 1 leads where y does, except when y is all ones below its
leading one (zero included), where |x| leads one place higher. That
corner is detected in parallel with the LZC and becomes a single
decrement. The negate still feeds the mantissa register, but it is off
the LZC path. LATENCY 6 keeps the original form, since it already has a
register between the two.

Verification:
- The identity is exhaustive over all 2^31 negative int32 values and
  holds on 200M int64 samples. The same harness flags the variant
  without the correction.
- A SIMULATION assertion compares the shift amount against the LZC of
  the magnitude. isa-32f-dsp passes with it, and fails on the known-bad
  mutation.
- Core rtlsim is cycle-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cache (302 MHz, LUT 28369, -5.8%) and vm (250.4 MHz, LUT 119695,
-12.4%; Fmax +9.2%) both improved past the gate tolerance with the
preceding timing fixes. These baselines were recorded from the finished
gate builds with --resume --update-baseline, at the user's request. The
full 12-DUT run on this tree then passes against them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pinning tensor's implementation strategy (ebc67f7, ce9bb0c) changed
its config hash. The baseline recorded before that was never refreshed,
so every gate run since has reported tensor STALE. It is re-recorded
here from a finished run on this tree (250.5 MHz, LUT 239260), which is
within tolerance of the old numbers. The full 12-DUT gate now passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@tinebp
tinebp force-pushed the fix/verilator-prebuilt-parts branch from feb8965 to 6510044 Compare September 27, 2026 08:45
@tinebp
tinebp merged commit 63c381a into master Sep 27, 2026
4 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant