This sample demonstrates how to implement an inference shader using some of the low-level building blocks from RTXNS. The sample loads a trained network from a file and uses the network to approximate a Disney BRDF shader. The sample is interactive; the light source can be rotated and various material parameters can be modified at runtime.
When the executable is built and run, the output shows a lit sphere using the neural network to approximate a Disney BRDF shader.
To load an inference neural network with RTXNS, several stages are needed which will be described in more detail below.
-
Create the host side neural network storage and initialize it
-
Create a GPU copy of the storage and initialize it with the host copy
-
Run the normal render loop calling the inference code instead of the disney shader to shade the sphere.
On the host, the setup and running of the neural network is quite simple and uses our RTXNS abstractions that are layered on top of the Graphics API Cooperative Vector extensions.
A rtxns::HostNetwork is created and initialized from a file. To ensure platform portability, the network should be stored in a non-GPU-optimized format, such as rtxns::MatrixLayout::RowMajor, and later converted to a GPU-optimized layout on the device, as shown below.
m_networkUtils = std::make_shared<rtxns::NetworkUtilities>(GetDevice());
rtxns::HostNetwork hostNetwork(m_networkUtils);
if (!hostNetwork.InitialiseFromFile(GetLocalPath("assets/data").string() + std::string("/disney.ns.bin")))
{
log::debug("Loaded Neural Shading Network from file failed.");
return false;
}
// Get a device optimized layout
rtxns::NetworkLayout deviceNetworkLayout = m_networkUtils->GetNewMatrixLayout(hostNetwork.GetNetworkLayout(), rtxns::MatrixLayout::InferencingOptimal);
This will load the network definition and parameters from the file, allocate a contiguous block of host memory for the parameters (weights and biases per layer), set the parameters from the file. At the same time we create a GPU optimal layout, rtxns::MatrixLayout::InferencingOptimal. Under the hood, this will use the CoopVector extensions to query the size of the allocations and perform the layout conversions.
Two parameter buffers are required, the first for the host layout and the second for device optimal layout. The parameter buffers contains all of the weights and biases for the network stored at a suitable precision, such as float16.
Once the host-layout buffer is populated, it can be converted to the device layout. The inference shaders use this buffer directly as input to Slang cooperative-vector functions.
// Create a buffer for the host side weight and bias parameters
nvrhi::BufferDesc bufferDesc;
bufferDesc.byteSize = hostNetwork.GetNetworkParams().size();
bufferDesc.debugName = "hostParamsBuffer";
bufferDesc.initialState = nvrhi::ResourceStates::CopyDest;
bufferDesc.keepInitialState = true;
m_mlpHostBuffer = GetDevice()->createBuffer(bufferDesc);
// Create a buffer for a device optimized parameters layout
bufferDesc.byteSize = deviceNetworkLayout.networkByteSize;
bufferDesc.canHaveRawViews = true;
bufferDesc.canHaveUAVs = true;
bufferDesc.debugName = "deviceParamBuffer";
bufferDesc.initialState = nvrhi::ResourceStates::UnorderedAccess;
m_mlpDeviceBuffer = GetDevice()->createBuffer(bufferDesc);
// Upload the parameters
const auto& params = hostNetwork.GetNetworkParams();
m_commandList->writeBuffer(m_mlpHostBuffer, params.data(), params.size());
// Convert to GPU optimized layout
m_networkUtils->ConvertWeights(hostNetwork.GetNetworkLayout(), deviceNetworkLayout, m_mlpHostBuffer, 0, m_mlpDeviceBuffer, 0, GetDevice(), m_commandList);
In this sample, the inference shader code is called directly from the pixel shader, so the render loop requires no modification apart from ensuring the parameter buffer is correctly bound.
As previously stated, this sample is designed to use a neural network to approximate the Disney BRDF shader. For reference, the shader code calling the Disney shader might look like this :
void main_ps(float3 i_norm, float3 i_view, out float4 o_color : SV_Target0)
{
//----------- Prepare input parameters
float3 view = normalize(i_view);
float3 norm = normalize(i_norm);
float3 h = normalize(-lightDir.xyz + view);
float NdotL = max(0.f, dot(norm, -lightDir.xyz));
float NdotV = max(0.f, dot(norm, view));
float NdotH = max(0.f, dot(norm, h));
float LdotH = max(0.f, dot(h, -lightDir.xyz));
//----------- Calculate core shader part DIRECTLY
float4 outParams = DisneyBRDF(NdotL, NdotV, NdotH, LdotH, roughness);
//----------- Calculate final color
float3 Cdlin = float3(pow(baseColor[0], 2.2), pow(baseColor[1], 2.2), pow(baseColor[2], 2.2));
float3 Cspec0 = lerp(specular * .08 * float3(1), Cdlin, metallic);
float3 brdfn = outParams.x * Cdlin * (1-metallic) + outParams.y*lerp(Cspec0, float3(1), outParams.z) + outParams.w;
float3 colorh = brdfn * float3(NdotL) * lightIntensity.rgb;
o_color = float4(colorh, 1.f);
}
We aim to swap out the DisneyBRDF() function to replace it with the neural network variant DisneyMLP. The rest of the code should not change.
The size of the neural network can be discovered from the file it is loaded from, but for simplicity it is also hardcoded into NetworkConfig.h which is shared by the application and shader code:
#define VECTOR_FORMAT half
#define TYPE_INTERPRETATION CoopVecComponentType::Float16
#define INPUT_FEATURES 5
#define INPUT_NEURONS (INPUT_FEATURES * 6) // Frequency encoding increases the input by 6 for each input
#define OUTPUT_NEURONS 4
#define HIDDEN_NEURONS 32
This network therefore contains 30 input neurons (5 input features encoded into 6 input parameters per feature), generates 4 output neurons and there are 32 neurons in each hidden layer. Each neuron supports float16 precision.
The Disney BRDF model used in this example encodes the inputs into the 0-1 frequency domain which is preferred by neural networks.
float params[INPUT_FEATURES] = { NdotL, NdotV, NdotH, LdotH, roughness };
inputParams = rtxns::EncodeFrequency<half, INPUT_FEATURES>(params);
The inference shader uses Slang's native CoopVec class to map the neural-network weights and biases to hardware in cooperative-vector form. See the Library Guide for more detail.
CoopVec<VECTOR_FORMAT, INPUT_NEURONS> inputParams;
The above code declares a native CoopVec type of size INPUT_NEURONS using precision format VECTOR_FORMAT. We know from the previous #defines that in this sample, these map to :
CoopVec<half, 30> inputParams;
There may be implementation-specific constraints on the underlying vector size, but Slang CoopVec objects can be declared with arbitrary sizes and the compiler pads them as required. Conceptually, these objects are similar to PyTorch tensors: each neural-network layer takes a cooperative vector as input and produces one as output. The input and output vectors may have different sizes.
To execute the inference model, rtxns functions use generic parameters to propagate cooperative vectors through the network. For example, LinearOp maps INPUT_NEURONS to OUTPUT_NEURONS using the VECTOR_FORMAT precision. This is equivalent to torch.nn.Linear in PyTorch.
hiddenParams = rtxns::LinearOp<
VECTOR_FORMAT, HIDDEN_NEURONS, INPUT_NEURONS,
CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(...)
The LinearOp function is a convenience wrapper in LinearOps.slang, built on the native CoopVec interface to perform a matrix multiply-add. The underlying function is shown below:
coopVecMatMulAdd<Type, Size>(...)
An activation function normally follows the linear operation. This example uses relu from the rtxns namespace in CooperativeVectorFunctions.slang.
hiddenParams = rtxns::relu(hiddenParams);
The linear regression and activation functions are called for each of the 4 layers of the network (1 input and 3 hidden). The final output is a float4 containing the result of the Disney approximation to be used when calculating the final color :
CoopVec<VECTOR_FORMAT, INPUT_NEURONS> inputParams;
CoopVec<VECTOR_FORMAT, HIDDEN_NEURONS> hiddenParams;
CoopVec<VECTOR_FORMAT, OUTPUT_NEURONS> outputParams;
// Encode input parameters, 5 inputs to 30 parameters
float params[INPUT_FEATURES] = { NdotL, NdotV, NdotH, LdotH, roughness };
inputParams = rtxns::EncodeFrequency<half, INPUT_FEATURES>(params);
// Forward propagation through the neural network
// Input to hidden layer, then apply activation function
hiddenParams = rtxns::LinearOp<VECTOR_FORMAT, HIDDEN_NEURONS, INPUT_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
inputParams, gMLPParams, weightOffsets[0], biasOffsets[0]);
hiddenParams = rtxns::relu(hiddenParams);
// Hidden layer to hidden layer, then apply activation function
hiddenParams = rtxns::LinearOp<VECTOR_FORMAT, HIDDEN_NEURONS, HIDDEN_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
hiddenParams, gMLPParams, weightOffsets[1], biasOffsets[1]);
hiddenParams = rtxns::relu(hiddenParams);
// Hidden layer to hidden layer, then apply activation function
hiddenParams = rtxns::LinearOp<VECTOR_FORMAT, HIDDEN_NEURONS, HIDDEN_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
hiddenParams, gMLPParams, weightOffsets[2], biasOffsets[2]);
hiddenParams = rtxns::relu(hiddenParams);
// Hidden layer to output layer, then apply final activation function
outputParams = rtxns::LinearOp<VECTOR_FORMAT, OUTPUT_NEURONS, HIDDEN_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
hiddenParams, gMLPParams, weightOffsets[3], biasOffsets[3]);
outputParams = exp(outputParams);
// Take the output from the neural network as the output color
return float4(outputParams[0], outputParams[1], outputParams[2], outputParams[3]);
Once the network has produced the 4 values, these are passed directly into the remainder of the shading code :
float3 Cdlin = float3(pow(baseColor[0], 2.2), pow(baseColor[1], 2.2), pow(baseColor[2], 2.2));
float3 Cspec0 = lerp(specular * .08 * float3(1), Cdlin, metallic);
float3 brdfn = outParams.x * Cdlin * (1-metallic) + outParams.y*lerp(Cspec0, float3(1), outParams.z) + outParams.w;
float3 colorh = brdfn * float3(NdotL) * lightIntensity.rgb;
o_color = float4(colorh, 1.f);
float4 DisneyMLP(float NdotL, float NdotV, float NdotH, float LdotH, float roughness)
{
uint4 weightOffsets = gConst.weightOffsets;
uint4 biasOffsets = gConst.biasOffsets;
CoopVec<VECTOR_FORMAT, INPUT_NEURONS> inputParams;
CoopVec<VECTOR_FORMAT, HIDDEN_NEURONS> hiddenParams;
CoopVec<VECTOR_FORMAT, OUTPUT_NEURONS> outputParams;
// Encode input parameters, 5 inputs to 30 parameters
float params[INPUT_FEATURES] = { NdotL, NdotV, NdotH, LdotH, roughness };
inputParams = rtxns::EncodeFrequency<half, INPUT_FEATURES>(params);
// Forward propagation through the neural network
// Input to hidden layer, then apply activation function
hiddenParams = rtxns::LinearOp<VECTOR_FORMAT, HIDDEN_NEURONS, INPUT_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
inputParams, gMLPParams, weightOffsets[0], biasOffsets[0]);
hiddenParams = rtxns::relu(hiddenParams);
// Hidden layer to hidden layer, then apply activation function
hiddenParams = rtxns::LinearOp<VECTOR_FORMAT, HIDDEN_NEURONS, HIDDEN_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
hiddenParams, gMLPParams, weightOffsets[1], biasOffsets[1]);
hiddenParams = rtxns::relu(hiddenParams);
// Hidden layer to hidden layer, then apply activation function
hiddenParams = rtxns::LinearOp<VECTOR_FORMAT, HIDDEN_NEURONS, HIDDEN_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
hiddenParams, gMLPParams, weightOffsets[2], biasOffsets[2]);
hiddenParams = rtxns::relu(hiddenParams);
// Hidden layer to output layer, then apply final activation function
outputParams = rtxns::LinearOp<VECTOR_FORMAT, OUTPUT_NEURONS, HIDDEN_NEURONS, CoopVecMatrixLayout::InferencingOptimal, TYPE_INTERPRETATION>(
hiddenParams, gMLPParams, weightOffsets[3], biasOffsets[3]);
outputParams = exp(outputParams);
// Take the output from the neural network as the output color
return float4(outputParams[0], outputParams[1], outputParams[2], outputParams[3]);
}
[shader("fragment")]
void main_ps(
VertexOut vOut,
out float4 o_color : SV_Target0)
{
float4 lightIntensity = gConst.lightIntensity;
float4 lightDir = gConst.lightDir;
float4 baseColor = gConst.baseColor;
float specular = gConst.specular;
float roughness = gConst.roughness;
float metallic = gConst.metallic;
// Prepare input parameters
float3 view = normalize(vOut.view);
float3 norm = normalize(vOut.norm);
float3 h = normalize(-lightDir.xyz + view);
float NdotL = max(0.f, dot(norm, -lightDir.xyz));
float NdotV = max(0.f, dot(norm, view));
float NdotH = max(0.f, dot(norm, h));
float LdotH = max(0.f, dot(h, -lightDir.xyz));
// Calculate approximated core shader part using MLP
float4 outParams = DisneyMLP(NdotL, NdotV, NdotH, LdotH, roughness);
// Calculate final color
float3 Cdlin = float3(pow(baseColor.r, 2.2), pow(baseColor.g, 2.2), pow(baseColor.b, 2.2));
float3 Cspec0 = lerp(specular * .08f * float3(1,1,1), Cdlin, metallic);
float3 brdfn = outParams.x * Cdlin * (1 - metallic) + outParams.y * lerp(Cspec0, float3(1), outParams.z) + outParams.w;
float3 colorh = brdfn * float3(NdotL) * lightIntensity.rgb;
o_color = float4(colorh, 1.f);
}
When compared with the original shader at the top of this section, it is clear that the only change is the original DisneyBRDF() function has been swapped out and replaced with the execution of the neural network in DisneyMLP().
