Procedural Terrain Rendering How-To

Well let me also share my progress so far, I have been doing planet rendering since quite a while, I did my first poorly written planet rendering algorithm in 2011 and many versions later, but this is the first time I was able to achieve a game ready performance of a planet renderer. “Its getting late here, so i will detail my algorithm hopefully in the morning or sometime when i am free in coming days” But here are some of the screenshots of the current work in progress.

Planar Terrain works perfect, though spherical version of the algorithm is currently a work in progress as there are some bugs in it. In a simple summary my algorithm is a combination of quadtree and toroidally updating clipmap grids with respect to the level of detail. Basically i use quadtree to generate terrain patch origins which are then compared with the origins of patches in a list of clipmap grid patches and rendered if the origins match. And using the camera position I update the clip map grid patches toroidally which allows me to reuse the same memory without requiring any new memory creation and deletion at runtime.

Let me also share some of the previous works with my planar terrain algorithm which basically runs purely on CPU, it runs at 60 fps but there are frame rate jumps. Though the current algorithms shown above all run on gpu.








I also played with some of the unity’s default image effects such as color grading and tonemapping, to enhance the look of the screenshots.

8 Likes

Beautiful stuff, yuneeb! I look forward to seeing more of your approach.

I tried to get closer to the problem of rendering multiple objects with DrawMesh. It seems to depend which “Dispatch()” call I do first in Start() to dispatch each computebuffer to the computeshader before it is sent to the vertex buffer. It seems the first ComputeBuffer.Dispatch() call overrules all following ones, as only the results from the first call are drawn (in my case, only the mesh of QuadtreeTerrain1).
Edit: To be more precise: Both meshes ares drawn but they seem to share the same locations and probably buffer.
I noticed that as the rendered triangles doubled with each Graphics.DrawMesh added.
Although from my point of unserstanding I am using different buffers in the separate QuadTrees, thus I’d expect that all are calculated separately in the computeshader and drawn separately in the OnRenderObject() function. I seem to be something missing when it comes to work with compute shaders and multiple buffers. Are there some kind of restrictions, e.g. that you can only once dispatch to a computeshader on the GPU within a limited time or frames, or something else?

yuneeb90: as already mentioned I really like your prototype. You CPU implementation looks nice with the different effects combined. Looking forward to where your GPU implementation progresses, and to share noise settings for the terrain etc.

1 Like

you need to intialize materials separately for every dispatch call. You may need to do it through script like this.

// Initialize a new material for every node / patch.

GridMaterial = new Material(GridShader);

// Update for every node / patch.

GridRendererCS.SetBuffer (0, "VertexBuffer", VertexCB);
GridRendererCS.SetBuffer (0, "NormalBuffer", NormalCB);
GridRendererCS.SetBuffer (0, "TexcoordBuffer", TexcoordCB);
GridRendererCS.SetBuffer (0, "TangentBuffer", TangentCB);
GridRendererCS.SetBuffer (0, "BinormalBuffer", BinormalCB);

GridRendererCS.Dispatch (0, 1, 1, 1);

GridMaterial.SetPass (0);
GridMaterial.SetBuffer ("VertexBuffer", VertexCB);
GridMaterial.SetBuffer ("NormalBuffer", NormalCB);
GridMaterial.SetBuffer ("TexcoordBuffer", TexcoordCB);
GridMaterial.SetBuffer ("TangentBuffer", TangentCB);
GridMaterial.SetBuffer ("BinormalBuffer", BinormalCB);

which means that you will not assign a new material through project view in unity rather you have to create a new material through script and to set material properties you will also need to do it through script for every material.

Thanky yuneeb90, appreciate the quick help,

EDIT [2015-11-25 08:33]: Guess I found it, as you made me try something out while looking at setting the buffer. Looks like you need to set the buffers again when you switch the kernel (sight :expressionless: ). After changing the Dispatch function, adding two lines to set the same buffer again after the first dispatch, it worked.

Seams betweens the patches appeared, need to look at the shader code again where the vertex positions are defined, but well, an important step forward.

//We then dispatch threads of our CSMain1 and CSMain2 kernel.
void Dispatch(QuadtreeTerrain quadtreeTerrain)
{
   // Set Buffers
   computeShader.SetBuffer(_kernel, "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
   computeShader.SetBuffer(_kernel, "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
   // Dispatch first kernel
   _kernel = computeShader.FindKernel("CSMain1");
   computeShader.Dispatch(_kernel, THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
   // Set Buffers
   computeShader.SetBuffer(_kernel, "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
   computeShader.SetBuffer(_kernel, "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
   // Dispatch second kernel
   _kernel = computeShader.FindKernel("CSMain2");
}

2 Likes

Writing from my phone and without access to the code, so forgive a short response. yuneeb is correct in that you need a separate material for each patch (for now, although I’ve submitted a feature request to Unity to allow MaterialPropertyBlock.SetBuffer() , which would allow you to reuse a single material and thus more efficient batching).

You should only need to search for the Kernels once, then save the results as a variable. Searching every time will have a significant performance impact.

And, just double checking: you should only have one ComputeShader object for each pass (and thus one kernel for each). Each patch would share these objects. Each patch needs its own material and own set of Buffers, of course.

Hi NavyFish, yuneeb,
followed your all your hints now (separate materials per patch, one compute shader object, moved the kernel search into the awake routine). The six patches now render quick and without problems.
EDIT 2:
Fixed the seam/space left between the patches.

I had to change the c# CPU code to

parameter.scale = 2.0f / (nVertsPerEdge);
parameter.spacing = 2.0f / (nVertsPerEdge-1.0f);

and the compute shader code to:

// First calculate the 'cube space' coordinates of the vertex:
float eastValue=  id.x - ((constants.nVertsPerEdge-1)/2.0); 
eastValue *= constants.spacing;                            
float3 cubeCoordEast = constants.cubeFaceEastDirection * eastValue;

// Do the same for the "north" direction:
float northValue=  id.y - ((constants.nVertsPerEdge-1)/2.0);
northValue *= constants.spacing;
float3 cubeCoordNorth = constants.cubeFaceNorthDirection * northValue;

I investigates further why my original problem occoured that not all planes were rendered, and found out it concentrated on the second kernel dispatch call. When I left the second dispatch call in, altough I did nothing yet in that second compute shader kernel, the patches did not render completely, especially the last one to be dispatched. When commenting the second kernel dispatch call out (which I do not need now but later for the normals), everything works withough any problems.

// Set Buffers
computeShader.SetBuffer(_kernel[0], "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
computeShader.SetBuffer(_kernel[0], "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
// Dispatch first kernel
computeShader.Dispatch(_kernel[0], THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
// Set Buffers
//computeShader.SetBuffer(_kernel[1], "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
//computeShader.SetBuffer(_kernel[1], "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
// Dispatch second kernel
//computeShader.Dispatch(_kernel[1], THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);

Makes me come to an important question:
Does the CPU stall and wait while the GPU works in the compute shader for the first dispatch? Or does it work async and the CPU immediately dispatches to the second kernel, no matter if the first kernel is done yet?
If the second is true, then I would have underestimated the orchestration of dispatches in the CPU and it would be clear why the above code wouldnt work.

Slowly catching up, fBm (Fractal Brownian Motion) added and coloring based on noise results. Right now I decide the coloring and so the terrain type completely in the surface shader. Will move this into the compute shader (decision of terrain type) once I figured out how to do the second compute shader kernel dispatch correct. Not anywhere as beautiful as your implementations, NavyFish and yuneeb90, but anyway happy to getting used to shaders more and more.

Of course I have yet no good indication of the performance, gain, however I am already impressed how fast the terrain renders even at very high octaves. Wow :open_mouth: . Looking forward to reimplement the quadtree LOD soon (should be more or less a copy-paste takeover of my old code fingers-crossed) and see how the performance behaves and continue playing with noise and increased terrain detail.

Going to extract the code to create a small Unity3D project for upload/link here now.
Would have liked to check how to do teh second dispatch to the compute kernel right, but anyway, besides that its pretty complete now to cover the whole topic initially, the small project should be a good starter fo the next one who needs to go the same route like me.

1 Like

@JoergZdarsky use an offset while generating vertices. I did it like this

x and z are indices in a single patch.

vx = (-HalfLodSize + EdgeLength * (x + 0.5f));
vz = (HalfLodSize - EdgeLength * (z + 0.5f));

this will perfectly align your grid patches. you will though need to see according to your algorithm how you will add that offset.

1 Like

The second is true.

All work on the GPU is done asynchronously from the CPU, but any command you send to the GPU is executed in the same sequence that you sent it. Sometimes half of the GPU will be working on one command and the other half on the next command. When the next command absolutely depends on the previous command having been completed, the graphics driver will ‘flush’ the GPU’s command queue up to that point, forcing it to catch up. This can manifest in GPU stalls.

But the CPU will rarely stall waiting for the GPU. This only happens when a resource must be synchronized between the two, such as during a CPU ‘read back’ of GPU data. In that case the GPU will flush, and the CPU will ‘block’, waiting for the GPU to finish and grant access to its resource.

Anyway, you should ‘ping pong’ or ‘double buffer’ your 2nd stage Compute Buffer data. In other words, don’t overwrite data in a buffer that was generated during the same frame. This will force the GPU to finish the first task before beginning the second task, thus destroying parallelism.

Another thing to consider is that switching between Kernels or Compute Shaders is expensive - so do all of your stage 1 processing first, and then switch to stage 2 and do all of that work, instead of swapping back and forth between stage 1 and stage 2 for each patch.

Another thing to look at in the Compute Shader’s code is StructuredBuffer vs RWStructuredBuffer. The former is read-only.

Hmm yes makes sense, however seems not so straighforward as I guessed. I right now try two write into a separate output-buffer “patchGenerationStageTwoDataBuffer” in kernel1, and read from that in kernel2 and write to the final “patchGeneratedDataBuffer” data buffer (which I then render in the shader).
Right now the mesh is not rendered, looks like kernel2 doesnt find its data in “patchGenerationStageTwoDataBuffer”.

#pragma kernel CSMain1
#pragma kernel CSMain2
#include "noiseSimplexOptimized.cginc"
 
//We define the size of a group in the x and y directions, z direction will just be one
#define threadsPerGroup_X 32
#define threadsPerGroup_Y 32

// The structure of the constants input computebuffer
struct GenerationConstantsStruct
{
    int nVertsPerEdge;
    float scale;
        float spacing;
        float3 patchCubeCenter;
    float3 cubeFaceEastDirection;
    float3 cubeFaceNorthDirection;
    float planetRadius;
    float terrainMaxHeight;
    float noiseSeaLevel;
    float noiseSnowLevel;
};

// The structure of the output computebuffer(s)
struct OutputStruct
{
        float4 position;
    float3 normal;
    float noise;
    float3 patchCenter;
};

//Various input buffers, and an output buffer that is written to by the kernel
StructuredBuffer<GenerationConstantsStruct>    generationConstantsBuffer;
RWStructuredBuffer<OutputStruct>        patchGenerationStageTwoDataBuffer;
RWStructuredBuffer<OutputStruct>        patchGeneratedDataBuffer;
 
//The kernel for this compute shader, each thread group contains a number of threads specified by numthreads(x,y,z)
//We lookup the the index into the flat array by using x + y * x_stride
//The position is calculated from the thread index and then the z component is shifted by the Wave function
[numthreads(threadsPerGroup_X,threadsPerGroup_Y,1)]

// Do position calculations
void CSMain1 (uint3 id : SV_DispatchThreadID)
{
    // Get the constants
    GenerationConstantsStruct constants = generationConstantsBuffer[0];
    float patchWidth = constants.spacing* constants.nVertsPerEdge;

    // Get outBuffOffset
    int outBuffOffset = id.x + id.y * constants.nVertsPerEdge;
    
    [ .. patch creation .. ]

    // Write to Output Buffer
    patchGenerationStageTwoDataBuffer[outBuffOffset].position = float4(patchCoordCentered.x, patchCoordCentered.y, patchCoordCentered.z, 1);
    patchGenerationStageTwoDataBuffer[outBuffOffset].normal = patchNormalizedCoord;
    patchGenerationStageTwoDataBuffer[outBuffOffset].noise = noise2;    
    patchGenerationStageTwoDataBuffer[outBuffOffset].patchCenter = patchCenter;

}

// Do normal calculations
void CSMain2 (uint3 id : SV_DispatchThreadID)
{
    // Get the constants
    GenerationConstantsStruct constants = generationConstantsBuffer[0];
    float patchWidth = constants.spacing* constants.nVertsPerEdge;

    // Get outBuffOffset
    int outBuffOffset = id.x + id.y * constants.nVertsPerEdge;

    [ .. do nothing, just write the data from the temporary buffer to the final buffer]

    // Read back from Input buffer to final output buffer
    patchGeneratedDataBuffer[outBuffOffset] = patchGenerationStageTwoDataBuffer[outBuffOffset];
}

I also tried to switch the buffers on the CPU (by setting the buffers new before the second dispatch) and make the buffer where kernel2 reads from readonly (StructuredBuffer), but doesnt seem to work either sight. Anyway will continue to try to get around that. Last topic left before adding the quadtree lod.

I will keep that in mind. Right now I only plan to have the normals calculation in the second stage, everything else should be ready in stage two. Only topic that comes to my mind when thinking about more stages would be to have some kind of fluid-simulation, but first lets form some nice terrain :yum:

EDIT 3 [2015-12-04]:
Got everything fixed now except the call of the second kernel for the normals calculation after the first kernel. I dont want to move on or focus on next topics until this is done. I currently bang my head against the wall while trying things out and checking the internet for information resources. ComputeShaders seem to be a topic with less information resource, even at the Unity docs, or I am missing something. There are few neat tutorials to get you started, but none of them deals with multiple stages in a ComputeShader. Strange, would have guessed this is an important use case. Anyway, its like it is.
Tried out various things to see how I can make sure kernel2 works on a completed buffer from kernel1. Writing into a separate buffer in kernel1 which kernel2 picks up before writing into another separate final buffer for rendering (like in my example above). Tried different orders of ComputeShader.SetBuffer() and ComputeShader.Dispatch() on the CPU. But right now, when I write like in the above code sample the vertex position results from one buffer into another in kernel2, nothing is rendered, which makes me think kernel2 picks up a buffer that is not yet completed in kernel2.
The only think that I can think of: in my code example the “condition” that a buffer is finished is based on each thread index (as you can see I simply write from buffer1 to buffer2 per index). Maybe that doesnt force the GPU to finish its task in kernel1 and I would need to somehow “swap” the whole buffer first somehow first before invoking the second dispatch. Not sure.
Anyone of you know some good resources in the net about using multiple buffers for multiple kernels, something that I might have missed?

Hey joerg… Didn’t get a notification of the new post, sorry to delay to long.

Can you try running the 2nd kernel during the next rendering frame? This should garunteed the first kernel finishes…

Well just a question out of curiosity, since @NavyFish @JoergZdarsky you guys are on this thread from the start, why is there a need for a multiple kernel in a compute shader for calculating planets vertex info? I was able to calculate almost everything in a single compute shader kernel or if you really want to compute something in a 2nd pass than the info that you passed on to the vertex/fragment shader from the compute shader you can use that shader as the second pass if you really need it.

Will give that a try. I can move the ComputeShader.Dispatch() calls into a couroutine that waits 1 frame inbetween. Should be simple, thank you for the hint, fingers crossed that this will work :smirk:

@yuneeb90: As we wrote per PM, I want to avoid additional noise calls (especially cellular noise was expensive in the past) and I hope that reusing the complete information (all vertice positions and noise values of a plane) in that output buffer from kernel1 in kernel2 will help me to avoid these noise calls (and just accessing the buffer with the right indices) when trying to get normals and slope information.

EDIT 2015-12-09: Did some try some tests with coroutines, still no success in skipping one frame, but now I got it down to some point that I realized that I am not able to dispatch anything to a second kernel at all. The second kernel simply is not being executed. I am only able to execute the first main routine CSMain1

    #pragma kernel CSMain1
    #pragma kernel CSMain2    
    void CSMain1 (uint3 id : SV_DispatchThreadID)
    {
       ....
    }
    void CSMain2 (uint3 id : SV_DispatchThreadID)
    {
      ....
    }

in/with the first kernel. But having a second kernel/function afterwards (e.g. CSMain2) it simply is not executed by the dispatch command in c#. Strange. Seems only the fist void function is being recognized by the dispatch command. But well, a point to investigate. But that means the timing of my kernel executions is not my problem - the problem of execution a second kernel is the issue.

EDIT 2015-12-09 (2): Done, both and especially the second stage / second kernel are now correctly invoked. So now I am ready to implement a two stage process in the compute shader to generate terrain. :heart_eyes:
The error was that in the compute shader I had both pragma definitions on top, afterwards both functions. Like:

#pragma kernel CSMain1
#pragma kernel CSMain2

[numthreads(threadsPerGroup_X,threadsPerGroup_Y,1)]

void CSMain1 (uint3 id : SV_DispatchThreadID)
{
    // code
}

void CSMain2 (uint3 id : SV_DispatchThreadID)
{
    // code
}

Things started to work when I put the code in different order:

#pragma kernel CSMain1

[numthreads(threadsPerGroup_X,threadsPerGroup_Y,1)]

void CSMain1 (uint3 id : SV_DispatchThreadID)
{
    // code
}

#pragma kernel CSMain2

[numthreads(threadsPerGroup_X,threadsPerGroup_Y,1)]

void CSMain2 (uint3 id : SV_DispatchThreadID)
{
    // code
}

There are a few (of the only few) compute shader tutorials around which describe my above initial implementation which didnt work. So anyone who has the same problem like me might try to change the order of codelines like above.
I verified this when adding in the second kernel an additional noise amount (0.8f) on top of the noise values of the first kernel’s noise result buffer . As expected, there is hardly any water remaining as I increased the height of the terrain in that second stage.

Now I can safely call this a complete example of invoking a compute shader in multiple stages and render its results in another shader, and extract a small project linked here for the next one to start. That weekend :grinning:
And then I can finally start with adding LOD by quadtree division. :yum:

Unless you use the derivative of the noise function, which limits your choice of function, can become extremely complex as your terrain function complexifies, and may have significant performance impacts, then the only way to determine normals for the generated terrain is to do so after all of the vertex positions have been generated. Thus you must do so in a second pass. Does that make sense?

I think thats a more intelligent solution. I am sort of like to get things working first and worry about optimizations later on. But doing normals on a separate pass is not only efficient but also makes some calculations easy. But the solution for the second pass I was suggesting was to use the vertex shader that receives compute shader data and calculating normals there, except that there is one problem that i am stuck at in passing the paramter.

uniform StructuredBuffer<float3> VertexBuffer;
uniform StructuredBuffer<float3> NormalBuffer;
uniform StructuredBuffer<float2> TexcoordBuffer;
uniform StructuredBuffer<float3> TangentBuffer;
uniform StructuredBuffer<float3> BinormalBuffer;

    v2f vert(uint id : SV_VertexID)
    { 				
         // I want to calculate normals in a 2nd pass here in the vertex shader, But I just tested out that the vertex
         // shader does not allow the following parameter "uint3 id : SV_DispatchThreadID" in place of 
         // uint id : SV_VertexID, Is there a way to pass the "uint3 id : SV_DispatchThreadID" parameter to the vertex
         // shader or some other solution?

    

         // **EDIT** Ok I got it working with the help of offsets, didnt changed the argument, I just added offset value to the 
         // "id" to get the respective vertex. But I think you will still need to generate an extra border vertex for normal
         //calculation, ( i am still working on it though) Note that my solution may not be the best, my code is a bit messed up at the moment due to experimentation.

 
 
        v2f OUT;
        OUT.Pos = VertexBuffer[id];
        OUT.vertex = mul(UNITY_MATRIX_MVP, float4(VertexBuffer[id], 1));
      	OUT.normal = NormalBuffer[id];
        OUT.texcoord = TexcoordBuffer[id];
        OUT.tangent = TangentBuffer[id];
        OUT.binormal = BinormalBuffer[id];
        return OUT;
    }

Plus I am stuck here in calculating per pixel normals, per vertex normals work fine for now (pics below), If anyone would like to share the solution to this problem, it would be much appreciated. thanks.

The texture can have more detail than just one point per vertex, so I don’t calculate my normals, tangents and bi-normals in the compute shader. Instead I generate a bump map at 2-3 levels deeper than the vertex data in the compute shader and then calculate the needed information from the bump map in the pixel shader.

Also those seams are probably there because when you calculate your normals and other data the adjacent sides of the planet are missing from the calculation.

1 Like

@cybercritic Thanks for the reply, yes you are right, the adjacent sides are missing from the calculation of planets cube patches. I was having a difficult time trying to figure this problem out, because since I am using skirts per patch to hide cracks, I thought the skirts may suffice for neighbor calculation even at the planet’s patch boundaries, but i am doing some wrong computations there.

Secondly your other solution related to calculating the bump map (per patch in a compute shader if i understood it right?) and calculating diffuse per pixel lighting from that. Well i will give that a shot I think it may solve all the major hurdles that i am facing right now. Thanks for the idea. This thread is getting quite popular on google searches, it will help spawn many planetary scale games in future :wink:

1 Like

I did this on my CPU version in the beginning also, at least for the texture (multiple resolution than the vertex density, e.g. vertex is 32x32 while the texture is 96x96). The noise calculations for the low density mesh and the higher density texture were separately done until I realized that I was throwing away a great chance for optimization. Because while doing lets say a 64x64 texture, I could reuse all the noise results from the 32x32 mesh result and only calculate the missing coordinates inbetween. So in the end I could skip half of the noise computations for the texture, which was a good win when it came to cellular noise. Although you can do this kind of separation in two different shaders too. It was in one of the first tutorials I’ve looked upon on procedural planets (need to check which one it was) that you dont want to throw away your noise results never ever. Having the same noise computation twice gives me a bad feeling.

Very intesting discussion, really! I will try to stick for two compute shader stages once I get this to work, as I find this more structured (I think all the terrain specifications should be done in the compute shader and I want the noise code for the terrain generation only once but not in two different shaders) and keep the vertex/surface shader more or less dumb. But this is maybe a kind of design decision or of what kind of code and structure you prefer. Looking forward to everyones progress!

This thread is sexy. And dangerous. I’ve got too much else to do and y’all are making me want to play around too!