Repository navigation
Close the fpga_gate and asic_gate: core/vm timing, dxa drain, baseline refresh - #419
Merged
Merged
Conversation
The shared-memory floor is the capacity a core keeps when bank count alone would leave it too small. At 8192 bytes a default NT=4 core holds only three of the four co-resident CTAs that `sgemm2_dxa_mcast` rendezvouses on, so its group barrier can never be satisfied and the launch makes no progress: both SimX and RTL hang, with no diagnostic from either. The floor now holds a whole warp-count group. Configurations at NT>=8 already derive more capacity from `LMEM_NUM_BANKS * lmem_bytes_per_bank` and are unchanged; NT<=4 regains the window it needs. rtlsim reproduces the recorded baseline for the case at 14344887 cycles against 14344871, and the full model_parity cell passes 29/29 with the two tracked known-issue cases behaving as annotated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The word-granular gather assembled a whole LMEM word per beat: a 512-bit window into the fill buffer, a six-stage spread network placing the run at the element stride, and a second six-stage barrel shifting it to the in-word offset. That is roughly 6.9k LUT of mux fabric in a unit that totals 10k, and twelve mux levels land in one path whose source is the drain-address recurrence, so it cannot be pipelined without giving back the throughput it exists to provide. The cost broke both synthesis gates: the Xilinx DUT fell to 249.3 MHz against a 300 MHz target (-17.7% on its baseline) at 14300 LUT (+41.6%), and the ASAP7 DUT grew to 133033 cells (+14%). The ASIC growth is the RTL alone -- rebuilding at the baseline's own toolchain reproduces it to within 1 cell -- so neither figure was a stale-baseline artifact. Draining one element per beat restores 305.4 MHz at 10260 LUT (+2%) on Xilinx, and 118381 cells (+2%) on ASAP7. The multicast and warp-group parity cells regain the known-issue marks that the wider drain had retired. This gives back a measured ~18% on the DXA-fed WGMMA shapes. Two ways to keep the gather were tried and measured first: folding the in-word offset into the read pointer, which cancels out to the same mux depth, and registering the write at the worker boundary, which bought 10 MHz for 631 LUT -- worse on the metric that is gated. The three WGMMA+DXA perf baselines still describe the wider drain and need regenerating; they are left untouched rather than edited, since the workload has moved since they were recorded. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The dxa entry was the one asic_gate build left on Yosys 0.57+157 / OpenSTA 2.7.0 when the rest moved to 0.69 / 3.1.0, so it read STALE on every nightly and its verdict said nothing about the design. Recorded at 1068.6 MHz, 12449 um^2, 118381 cells. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ack AGU The core fpga_gate DUT (NT16/NW16, xcu55c) had fallen from 298.1 MHz to 272.8 MHz. Aggregating the violating paths of every failing build put 183 of 300 in VX_scoreboard, the rest in the LSU request path; two causes, both latency-neutral to fix. Scoreboard: operands_ready_n was computed after the issue decision, so operands_ready_r -> 16-way out_arb grant -> staging_fire -> reserve rd and pick the next instr's mask -> NUM_REGS-wide reduce -> operands_ready_r all fit in one cycle. Readiness is now evaluated for both outcomes (staging holds / staging fires) from grant-independent state -- one reduction per candidate instr plus a register compare for the reserved rd -- and the grant only selects. Commit also registers a per-warp one-hot eop_wis ahead of its output buffer, so a warp's release reads its own flop instead of every warp decoding one fanout-52 wis. Stack AGU: the stack interleave put a subtract, a permute and an add, three dependent full-width CPAs plus two full-width compares, between the LSU lane dispatch and the request queue. The permute stays inside the group field, so the base cancels on the high bits; what remains is a narrow correction the width of the group field above the window base's trailing zeros. Plain and pack-load forms share one adder, and the window test compares only bits above the stack size. Verified: - fpga_gate core, at the tip of this series: 300.8 MHz, 178,623 LUT (+4.0%) -- PASS. top, tensor and gfx unchanged within tolerance. - core rtlsim: 25/25 runs cycle- and instruction-identical to before, 9/9 model_parity pass. No SimX change: no cycle moved. - AGU equivalence vs the direct remap: 168 configs, 6.5M addresses, including non-power-of-two core counts; 0 mismatches. - hw/unittest issue, mem_scheduler, generic_queue, core: clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The vm fpga_gate build missed its 237.5 MHz floor (234 MHz). Its worst paths all ended at the L2 bank's core response queue, the dmmu output stage and the L2 fill crossbar. Every change below is cycle-exact. L2 AMO engine: - With a deferred commit stage, the write-back queue's lane hits are registered from the next commit key, folding in slot occupancy, so the forward network and the response start from flops. - Coalesce and slot selection are computed for both drain outcomes, so the late wb_fire only picks between them. - A queued result drops its bytes from the settling entry. Each byte then has at most one writer, and the forward merge is an OR-tree instead of a priority mux. - The in-place AMO response is the bank's forwarded word under the byte mask, which takes the lane select off the response path. - The bank's S0 atomic hit prediction is the raw tag match. Atomics are exempt from the hit-order hazard, and this keeps the MSHR probe off the commit_busy cone. dmmu: - The TLB permission check runs per entry beside the tag compare, rather than behind the selected entry's flags. - The park lane is a one-hot select of the lane categories. L2 fill crossbar: OUT_BUF 2 registers the payload at the input handshake, so it no longer loads on the bank's late ready. Latency is unchanged. Results: - fpga_gate, at the tip of this series: vm 250.0 MHz (was 234.2), LUT 119,682 (-12.4%). top 250.0 MHz, LUT +4.5% (was +5.4%, over the cap); the dedup of the forwarded word in the AMO engine accounts for that recovery. cache 301.8 MHz, LUT 28,369 (-5.8%). - rtlsim amo+vm: 31/31, cycle- and instruction-identical to the pre-change run; amo base/threads8 and vm tlb-stress-l1 re-run at the series tip. - The SIMULATION checks on the registered hits and on the settling-entry disjointness fire on known-bad mutations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The core fpga_gate build fell below its 283.2 MHz floor (268.5 MHz). The worst paths ran from SFU lane dispatch through the 12-bit CSR address decode, the scheduler's per-warp trap-CSR muxes and the lane replication, into the CSR response buffer and back into the scheduler's trap registers. Every CSR request is already held one cycle so the CTA context RAM can be read. That cycle now does the decode. Read and write path: - The stable value, a one-hot of the sources that can still move (counters, thread/warp masks, fcsr, the CTA RAM), the write target and the read-modify-write result are all registered in the wait cycle. - The fire cycle is a small AND-OR, and the trap-CSR write is a one-hot target, so the scheduler compares no addresses. - Per-lane thread ids are the lane index on top of a registered base. - A SIMULATION check compares every registered value against its live form at fire. Host counter reads get their own MPM-window decode and no longer wait on a CSR request. The kernel-path decode drops the MPM tree, which always read as zero for class 0. Two latent bugs are fixed along the way: - CSR writes are qualified by fire, not by the request being valid. With the response buffer full, a CSRRW used to write early and then return the new value. - VX_dcr_data latches the MPM class with its address, instead of reading it live off the DCR bus after the request has left. Results: - fpga_gate: core 286 MHz (floor 283.2). The CSR paths left the worst 100; SFU lane dispatch shrinks by about 300 LUTs per core. - rtlsim core: all 34 perf lines are cycle- and instruction-identical to the pre-change run. The 16 failures there are the pre-existing stale perf-baseline config hashes, the same set before and after. - The capture assertion fires on a known-bad mutation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After the CSR fix, core's worst path ran from the dispatcher into fcvt. At LATENCY 5 (XLEN=32), the integer negate and the LZC share a cycle, fed across a long route from the dispatcher's output register. The LZC now reads the one's complement y = ~x of a negative integer. |x| = y + 1 leads where y does, except when y is all ones below its leading one (zero included), where |x| leads one place higher. That corner is detected in parallel with the LZC and becomes a single decrement. The negate still feeds the mantissa register, but it is off the LZC path. LATENCY 6 keeps the original form, since it already has a register between the two. Verification: - The identity is exhaustive over all 2^31 negative int32 values and holds on 200M int64 samples. The same harness flags the variant without the correction. - A SIMULATION assertion compares the shift amount against the LZC of the magnitude. isa-32f-dsp passes with it, and fails on the known-bad mutation. - Core rtlsim is cycle-identical. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cache (302 MHz, LUT 28369, -5.8%) and vm (250.4 MHz, LUT 119695, -12.4%; Fmax +9.2%) both improved past the gate tolerance with the preceding timing fixes. These baselines were recorded from the finished gate builds with --resume --update-baseline, at the user's request. The full 12-DUT run on this tree then passes against them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pinning tensor's implementation strategy (ebc67f7, ce9bb0c) changed its config hash. The baseline recorded before that was never refreshed, so every gate run since has reported tensor STALE. It is re-recorded here from a finished run on this tree (250.5 MHz, LUT 239260), which is within tolerance of the old numbers. The full 12-DUT gate now passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tinebp
force-pushed
the
fix/verilator-prebuilt-parts
branch
from
September 27, 2026 08:45
feb8965 to
6510044
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This branch makes both synthesis gates pass on every DUT. Each commit fixes one root cause, and none of them moves a cycle count.
Commits
fe67ae44bsgemm2_dxa_mcastwaits on, so SimX and RTL both hung.26f4882d6727b8710ba7888b31487beeb8b6LATE_VALIDskid mode inVX_stream_buffer. L2 fill crossbar:OUT_BUF2.d5a47950eVX_dcr_datalatches the MPM class with its address.428d65c11~xfor a negative integer, plus a power-of-two correction, instead of waiting on the negate.248fb33b6feb8965efimpl_strategy(ebc67f7a9,ce9bb0c55) changed its config hash without a refresh, so it had read STALE since 08-28.Verification
fpga_gate: full 12-DUT run on the final tree, all PASS:
asic_gate: full 8-DUT run, all PASS, every DUT within ±1% of its baseline.
rtlsim, cycle- and instruction-identical to pre-change runs:
isa-32f-dsp, the only config that exercises the RTL fcvt. The default rtlsim FPU is DPI.New SIMULATION assertions: each new registered or look-ahead path is checked against the logic it replaces. Each check was confirmed to fire on a known-bad mutation of its own path:
Libraries: the
VX_stream_bufferLATE_VALIDmode is off by default. It matched the accept-enabled form over 4M random cycles, and a known-bad variant fails.No SimX change: no cycle count moved. That is covered by the rtlsim identity above, plus model_parity for the config and core commits.
Known, not addressed here
perf_gateonly because their perf-baseline config hashes are stale. The failing set is identical before and after this branch, and regenerating them is a separate reviewed step. The three WGMMA+DXA perf baselines (see26f4882d6) are in the same state.csrw mepc; mretandcsrw mtvec; ecallare not ordered, because CSR ops are notwstall. No current test reaches it, including therv32ui-pstartup, which does exactlycsrw mepc; mret. It needs a directed test and a matching SimX decode change.build32_asic_gatemust be run withTOOLDIR=~/tools; its configured TOOLDIR points at a path that no longer exists.🤖 Generated with Claude Code