A Verilog implementation of an 8-bit SPI master/slave subsystem connected to a 256-byte RAM. The design is organized as a complete memory communication path, with a host-facing SPI master, an internal SPI bus, a slave controller, and a RAM subsystem contained inside a single top-level wrapper.
SPI stands for Serial Peripheral Interface. It is a synchronous serial communication protocol commonly used for short-distance communication between digital devices such as microcontrollers, FPGAs, memories, sensors, converters, and displays.
A conventional SPI bus uses four signals:
SCLK: serial clock generated by the master.MOSI: Master Out, Slave In. Carries data from the master to the slave.MISO: Master In, Slave Out. Carries data from the slave to the master.SS_n: active-low slave-select signal that frames a transaction.
Because MOSI and MISO are separate data lines, SPI supports full-duplex communication. The clock is supplied by the master, so both sides exchange data relative to defined SCLK edges rather than using an independent baud-rate recovery mechanism.
This project applies SPI to a small FPGA memory subsystem. The current design uses an 8-bit, multi-byte transaction format. A command byte identifies the requested RAM operation, and a second byte carries the corresponding address, write data, or dummy read payload. The master can keep SS_n asserted across both bytes so the slave treats them as one continuous transaction.
The spi_wrapper module exposes a compact parallel control interface to the host system.
| Signal | Direction | Width | Description |
|---|---|---|---|
clk |
Input | 1 | Main system clock. The provided constraints define a 100 MHz clock. |
rst_n |
Input | 1 | Active-low asynchronous reset. |
master_start_tx |
Input | 1 | Starts transmission of one SPI byte. |
master_hold_ss |
Input | 1 | Keeps SS_n low after the current byte when set to 1. |
master_tx_data |
Input | 8 | Byte presented to the SPI master for transmission. |
master_tx_ready |
Output | 1 | Pulses when the current byte transfer completes. |
master_rx_data |
Output | 8 | Byte received by the master from MISO. |
busy |
Output | 1 | Indicates read-response activity inside the slave path. |
The master and slave are connected internally inside spi_wrapper.
| Signal | Source | Destination | Function |
|---|---|---|---|
SCLK |
SPI master | SPI slave | Serial transfer clock. |
SS_n |
SPI master | SPI slave | Active-low transaction select. |
MOSI |
SPI master | SPI slave | Serial command and payload data. |
MISO |
SPI slave | SPI master | Serial readback data. |
Each memory operation is formed from two consecutive 8-bit SPI bytes while SS_n remains low.
| Byte | Purpose | Description |
|---|---|---|
| Byte 1 | Command phase | Selects the RAM operation. |
| Byte 2 | Payload phase | Carries an address, write data, or a dummy value for a read. |
The host sends the command byte with master_hold_ss = 1, then sends the payload byte with master_hold_ss = 0 to finish the transaction.
| Command | Operation | Payload byte | Result |
|---|---|---|---|
0x00 |
Set write address | 8-bit address | Updates the internal write pointer. |
0x01 |
Write data | 8-bit data | Writes the payload to the stored write address. |
0x02 |
Set read address | 8-bit address | Updates the internal read pointer. |
0x03 |
Read data | Dummy byte, typically 0x00 |
Returns the stored byte on MISO during the payload transfer. |
The SPI transfer follows Mode 0 behavior in the current RTL:
- SCLK is idle low.
- MOSI is sampled by the slave on rising SCLK edges.
- MISO advances on falling SCLK edges and is sampled by the master on the following rising edge.
- All main modules use an active-low asynchronous reset.
The SPI master defines CLK_DIV = 4. In the current implementation, SCLK toggles once every four system-clock cycles, so a complete SCLK period takes eight system-clock cycles. With a 100 MHz clk, the resulting SCLK frequency is 12.5 MHz.
The system specification describes the default SCLK as system clk / 4. The current RTL therefore differs from that stated frequency relationship and should be treated as system clk / 8 unless the divider logic is revised.
flowchart LR
H[Host interface] -->|start, hold_ss, tx_data| M[SPI master]
M -->|MOSI, SCLK, SS_n| S[SPI slave]
S -->|rx_cmd, rx_data, valid strobes| R[256 x 8 RAM]
R -->|tx_data, tx_valid| S
S -->|MISO| M
M -->|tx_ready, rx_data| H
The design is split into four Verilog modules.
| Module | Role |
|---|---|
spi_wrapper.v |
Instantiates the master, slave, and RAM and connects the internal SPI and control signals. |
spi_master.v |
Generates SCLK and SS_n, serializes outgoing bytes, receives MISO data, and supports continuous two-byte transfers through hold_ss. |
spi_slave.v |
Receives SPI bytes, separates command and payload phases, generates RAM control strobes, and serializes RAM read data onto MISO. |
ram_sp_async.v |
Implements a 256 x 8 memory array with stored read and write addresses and command-based access. |
spi_master uses three states: IDLE, TX, and DONE. A transfer begins when start_tx is asserted. The master lowers SS_n, shifts the most-significant bit first on MOSI, generates SCLK from the system clock, and captures MISO on rising SCLK edges.
At the end of each byte, tx_ready is asserted and the received byte is copied to rx_data. If hold_ss is high, SS_n remains low so the next byte continues the same SPI frame. If hold_ss is low, SS_n returns high and the transaction ends.
spi_slave uses an idle state plus two active transaction phases: CMD_PHASE and DATA_PHASE. SCLK is synchronized into the system-clock domain with a three-bit shift register, and rising and falling edges are detected from the synchronized samples.
During CMD_PHASE, the slave collects eight MOSI bits into rx_cmd and pulses cmd_valid. During DATA_PHASE, it collects the next eight bits into rx_data and pulses rx_valid.
For command 0x03, the early cmd_valid pulse allows the RAM to prepare read data before the second byte is fully clocked. When tx_valid arrives, the slave loads the RAM output into its transmit shift register and returns the byte through MISO.
The RAM contains 256 locations, each eight bits wide. It keeps separate write and read address registers.
0x00updateswr_addrfromrx_data.0x01storesrx_dataatmem[wr_addr].0x02updatesrd_addrfromrx_data.0x03placesmem[rd_addr]ondoutand pulsestx_validwhencmd_validis observed.
Although the module is named ram_sp_async, the current RTL performs address updates, writes, and read-response generation inside a clocked always block.
A write to RAM uses two SPI transactions:
- Send
0x00, then the target address. - Send
0x01, then the data byte.
A read uses two SPI transactions:
- Send
0x02, then the target address. - Send
0x03, then a dummy byte such as0x00. The requested RAM byte is returned over MISO during this second byte.
Verification is driven by tb_spi_wrapper.v and the ModelSim script run.do. The testbench generates a 100 MHz clock, applies reset, and uses a reusable send_byte task to exercise the host-facing master interface.
The current directed test performs the following sequence:
- Set the write address to
0x05. - Write
0xAAto that address. - Set the read address to
0x05. - Issue a read command followed by a
0x00dummy byte. - Print the received value and compare it visually with the expected value
0xAA.
run.do compiles the four RTL modules and the testbench, launches tb_spi_wrapper, adds the main system and SPI signals to the waveform window, and runs the simulation.
The saved ModelSim transcript shows clean RTL compilation for all modules, with zero compiler errors and zero compiler warnings. Two runtime warnings are related to the waveform log file already being in use and do not indicate RTL compilation problems.
The functional readback is not currently passing in the saved simulation. The transcript reports:
Final Read Data Output: 0xxx (Expected: 0xaa)
For that reason, the current repository should be described as compiling successfully but still requiring debug on the read-data path before functional verification can be considered complete.
The testbench is directed rather than self-checking. It prints the expected and observed value but does not use an assertion or pass/fail counter.
The included waveforms show the command and payload transfers together with SS_n, SCLK, MOSI, MISO, master_tx_ready, and master_rx_data.
The provided XDC targets the Xilinx Artix-7 xc7a200tfbg484-3 device and defines the system timing assumptions used for implementation.
| Constraint | Value |
|---|---|
| System clock | 100 MHz |
| Clock period | 10.0 ns |
| Clock uncertainty | 0.20 ns |
| Maximum input delay | 1.00 ns |
| Minimum input delay | 0.50 ns |
| Maximum output delay | 1.50 ns |
| Minimum output delay | 0.50 ns |
The reset is declared as a false timing path. Input delays are applied to master_start_tx and master_tx_data[*]. Output delays are applied to master_tx_ready, busy, and master_rx_data[*].
The current XDC does not apply input-delay constraints to master_hold_ss, and it does not include package-pin assignments or I/O-standard declarations. These items should be completed before using the wrapper as a board-level top module.
The SPI signals themselves are internal nets between the master and slave, so the current wrapper does not expose MOSI, MISO, SCLK, or SS_n as FPGA package pins.
The repository includes Vivado screenshots for utilization, timing, placement, schematic, and power analysis. These reports were captured from an earlier RTL revision and should not be treated as current sign-off results for the updated source files.
The saved schematic and timing-path screenshots show 10-bit master_tx_data and master_rx_data buses, while the current RTL uses 8-bit buses and adds the master_hold_ss input. The reports should therefore be regenerated after the current RTL is functionally corrected and resynthesized.
The saved hierarchical utilization report contains the following values for spi_wrapper:
| Resource | Used | Available | Utilization |
|---|---|---|---|
| Slice LUTs | 744 | 133,800 | 0.56% |
| Slice registers | 2,195 | 267,600 | 0.82% |
| F7 muxes | 272 | 66,900 | 0.41% |
| F8 muxes | 136 | 33,450 | 0.41% |
| Slices | 813 | 33,450 | 2.43% |
| LUTs as logic | 744 | 133,800 | 0.56% |
| Bonded I/O | 25 | 285 | 8.77% |
| BUFGCTRL | 1 | 32 | 3.12% |
The archived hierarchy assigns most of the logic to the RAM block, with 689 LUTs and 2,100 registers reported for u_ram_sp_async.
The saved Vivado timing summary reports all constraints met for that earlier implementation.
| Metric | Archived result |
|---|---|
| Worst negative slack, WNS | 0.436 ns |
| Total negative slack, TNS | 0.000 ns |
| Setup failing endpoints | 0 |
| Worst hold slack, WHS | 0.064 ns |
| Total hold slack, THS | 0.000 ns |
| Hold failing endpoints | 0 |
| Worst pulse-width slack, WPWS | 4.500 ns |
| Total pulse-width negative slack, TPWS | 0.000 ns |
| Pulse-width failing endpoints | 0 |
The saved device view shows the placed logic occupying a small region of the Artix-7 fabric in the earlier implementation.
The saved power report estimates the following values for the earlier implementation:
| Metric | Archived result |
|---|---|
| Total on-chip power | 0.137 W |
| Dynamic power | 0.006 W |
| Device static power | 0.131 W |
| Estimated junction temperature | 25.3 C |
| Confidence level | Low |
Because Vivado marks the confidence level as low and the report belongs to an older RTL revision, the values should be treated as reference estimates rather than measurements for the current design.
| Contributor |
|---|
| Abdelrahman Ehab |
| Omar Walid |
| Nourhan Hussain |







