Procedural Terrain Rendering How-To

EDIT: Post shortened, tldr :wink:

Hi NavyFish,

Can you please detail these sections?
What I right now do is:

  • Create a ComputeBuffer “verticePosBuffer” with 224x224 vertice-positions (input for this is the flat prototype plane)

  • Create a ComputeBuffer “trianglesBuffer” with the triangle indices (input for this is the flat prototype plane)

  • Create a ComputeBuffer “noiseBuffer” in which 224x224 float values are written into in the Compute Shader

    void CreateBuffers(){
    int VERTICESCOUNT = this.plane.vertices_pos.Length;
    int TRIANGLESCOUNT = this.plane.triangles.Length;
    verticesPosBuffer = new ComputeBuffer(VERTICESCOUNT, 12); // (Vector3 -> 12 bytes)
    verticesPosBuffer.SetData(this.plane.vertices_pos);
    trianglesBuffer = new ComputeBuffer(TRIANGLESCOUNT, sizeof(int));
    trianglesBuffer.SetData(this.plane.triangles);
    noiseBuffer = new ComputeBuffer(VERTICESCOUNT, sizeof(float));
    outputBuffer = new ComputeBuffer(VERTICESCOUNT, 12); //Vector3 -> 12 bytes)
    }

  • Pass all three buffers to a Compute Shader that
    – fills the “noiseBuffer” with a heightmap (224x224 floats)
    – displaces the vertices from the verticesPosBuffer and fills the outputBuffer with the displaced vertices.

    void Dispatch() {
    computeShader.SetBuffer(_kernel, “vertPosBuff”, verticesPosBuffer);
    computeShader.SetBuffer(_kernel, “noiseBuff”, noiseBuffer);
    computeShader.SetBuffer(_kernel, “output”, outputBuffer);
    computeShader.Dispatch(_kernel, 32, 32, 1);
    }

When this is done I pass the buffers to the render shader:

void OnRenderObject()  {
   int VERTICESCOUNT = plane.vertices_pos.Length;
   int TRIANGLESCOUNT = plane.triangles.Length;
   Dispatch();
   material.SetPass(0);
   material.SetBuffer("buf_vertices_pos", outputBuffer);
   material.SetBuffer("buf_triangles", trianglesBuffer);
   material.SetBuffer("buf_noise", noiseBuffer);
   Graphics.DrawProcedural(MeshTopology.Triangles, TRIANGLESCOUNT);
}

I understood that you don’t pass the vertice-positions to the compute shader but only the boundary information (width of plane etc.) to create the noiseBuffer. Makes sense, a switch to that should be straightforward. But I have a problem to understand how then proceed in rendering multiple planes out of one initial prototype plane from the CPU and how you pass positions to a vertex shader.

Do you ever create a verticesPositionBuffer like me and pass it over (once or many times for each plane) to a vertex shader?
Which information (buffers) do you pass to the vertex shader?
You said you create a mesh as workaround. Do you create for each plane a UnityEngine.GameObject or UnityEngine.Mesh to apply a stock surface shader? And if so, how do you displace the vertices with a custom vertex shader while you want to use a surface shader the same time?
I think seeing how you’d do the section “OnRenderObject()” would be very interesting.

Thanks a lot beforehand!

PS: Pretty excited to get plane creation on the GPU running and then combine this with the current implementation. This parallel processing of vertices (I did 32x32 planes before in serial on the CPU) is really impressive and along with the shared vertices and leaving everything on the GPU should give a massive performance boost. Also, when switching to 224x224 planes I might skip a few highest levels of my quadtree as there is already enough precision on lower nodes.

Can’t wait to give a detailed response. I’ve been out of town all weekend and am traveling home today.

Great job figuring out how to use Compute Buffers in Unity! I’ll have to refer to my code to address a few details but you’re definitely on the right path.

You’re correct that I don’t pass a base plane or patch to the Compute Shader’s - just basic information relevant to that patch’s generation, such as size and position on the planet. And correct regarding the use of Compute Buffers, but the only ones I fill are position, normals, and (eventually, but not in my current implementation) texture coordinates. Position represents the final vertex positions, ie scaled appropriately for the LOD, curved to match the surface of a sphere, and petrurbed by a noise function.

The mesh (not procedural mesh) I use already has its triangles array defined. You only ever need to do this once. I draw the mesh normally, but give it a custom surface shader, which has a custom ‘per vertex pass’ (forget the actual name at the moment, but one of the Unity Shader tutorials demonstrates this - its the tutorial that displaces vertices of a model face before drawing them). This vertex pass reads from the positions Compute Buffer and modifies the position of the ‘output’ vertex based upon the position read from the input Compute Buffer (used as a StrucutredBuffer in the surface Shader). If it weren’t for this vertex pre-pass,the plat plane would be drawn.

I will show you my code as soon as I’m back at a computer. Good work!

edit added 2nd to last paragraph

1 Like

Okay, back home. i’m going to post my code and provide some explanation ws we go, but don’t have a ton of time at the moment to dive deep into it. Hopefully with the progress you’ve made so far it should be fairly self-explanatory. I’ve tried to remove as much of the ‘non-relevant’ code as possible. Also note that I’m in the midst of a somewhat poor refactoring job, attempting to ‘modularize’ certain aspects of the patch generation. I somehow failed to make a backup of the code prior to this :confused: so it’s not in a compilable state at the moment. That shouldn’t matter - the core code that drives Unity is unchanged, but the organization of it is going to be wrong, so if you see random custom classes referenced in some places but not others, don’t think much of it.

First off, creation of the prototype patch/mesh (dummy mesh). I do this just once, in a class “PatchManager” that is a ‘singleton’ MonoBehaviour.

    private void setupDummyMesh()
    {
        int nVerts = structuralConfiguration.nVerts;
        int nVertsPerEdge = structuralConfiguration.nVertsPerEdge;
        Vector3[] dummyVerts = new Vector3[nVerts];
        Vector2[] uv0 = new Vector2[nVerts];
        int[] triangles = new int[(nVertsPerEdge - 1) * (nVertsPerEdge - 1) * 2 * 3];

        float height = 0;
        for (int r = 0; r < nVertsPerEdge; r++)
        {
            int rowStartID = r * nVertsPerEdge;
            for (int c = 0; c < nVertsPerEdge; c++)
            {
                int vertID = rowStartID + c;
                dummyVerts[vertID] = new Vector3(c, height, r);
                Vector2 uv = new Vector2();
                uv.x = r / (float)(nVertsPerEdge - 1);
                uv.y = c / (float)(nVertsPerEdge - 1);
                uv0[vertID] = uv;
            }
        }

        int triangleIndex = 0;
        for (int r = 0; r < nVertsPerEdge - 1; r++)
        {
            int rowStartID = r * nVertsPerEdge;
            int rowAboveStartID = (r + 1) * nVertsPerEdge;
            for (int c = 0; c < nVertsPerEdge - 1; c++)
            {
                int vertID = rowStartID + c;
                int vertAboveID = rowAboveStartID + c;

                triangles[triangleIndex++] = vertID;
                triangles[triangleIndex++] = vertAboveID;
                triangles[triangleIndex++] = vertAboveID + 1;

                triangles[triangleIndex++] = vertID;
                triangles[triangleIndex++] = vertAboveID + 1;
                triangles[triangleIndex++] = vertID + 1;
            }
        }

        dummyMesh = new Mesh();
        dummyMesh.vertices = dummyVerts;
        dummyMesh.uv = uv0;
        dummyMesh.SetTriangles(triangles, 0);
    }

I also go ahead and set up the material which will be used for later rendering. This is also done only once, in a ‘singleton’ monobehavior. Note that the “ProceduralMeshVertSurf” file must be located in the project’s Resources folder.

    material = new Material(Shader.Find("ProceduralMeshVertSurf"));
    material.SetFloat("_Metallic", 0);
    material.SetFloat("_Glossiness", 0);
    Texture2D texture = (Texture2D)UnityEngine.Resources.Load("GrassRockyAlbedo");
    material.SetTexture("_MainTex", texture);

Now, for each patch I create two ComputeBuffers.

generationConstants, which is where I’ll put the “inputs” to the ComputeShader - i.e. properties required to build the patch like vertex spacing, patch world center, etc.

computeShader.setBuffer(kernel, "terrainGenerationConstants", generationConstants); 

and

patchGeneratedDataBuffer, which the ComputeShader will fill with vertex data (position, normal, etc).

computeShader.setBuffer(kernel, "patchOutput", patchGeneratedDataBuffer); 

Note that the strings above must match those found in the ComputeShader below (I actually do this procedurally but have used plain strings in this example for clarity).

After filling the generationConstants buffer with the appropriate data, I call ComputeShader.setBuffer for the above two buffers, and then dispatch the Compute Shader with

    public void dispatch()
    {
        computeShader.Dispatch(kernel,
        THREADGROUP_SIZE_X,
        THREADGROUP_SIZE_Y,
        THREADGROUP_SIZE_Z);
    }

These constants are defined as:

        public static int nVertsPerEdge { get { return 224; } }     //Should be multiple of 32
        public static int nVerts { get { return nVertsPerEdge * nVertsPerEdge; } }
        public int THREADS_PER_GROUP_X { get { return 32; } }
        public int THREADS_PER_GROUP_Y { get { return 32; } }
        public int THREADGROUP_SIZE_X { get { return nVertsPerEdge / THREADS_PER_GROUP_X; } }
        public int THREADGROUP_SIZE_Y { get { return nVertsPerEdge / THREADS_PER_GROUP_Y; } }
        public int THREADGROUP_SIZE_Z { get { return 1; } }

Here’s the basic version of Compute Shader itself. I’ve left the preprocessor stuff in just to demonstrate a cool technique, but it’s definitely not essential:

#pragma kernel CSMain

#define threadsPerGroup_X 32
#define threadsPerGroup_Y 32
#define nVerticesPerSide 224
#define nVerticesPerSideFloat 224.0

#define TWO_PI 6.283185

#include "noiseSimplex.cginc"
#include "2DNoiseFunctions.cginc"

//#define     RIDGID
#define   HYBRID

#ifdef HYBRID
//Good hybridMultifractal values
#define     NoiseFrequency              0.001
#define     OneMinusFractalIncrement    0.3
#define     Lacunarity                  1.918
#define     nOctaves                    14
#define     MultifractalOffset          0.9

#elif defined RIDGID
//Good ridgedMulti values
#define     NoiseFrequency              0.015
#define     OneMinusFractalIncrement    0.8
#define     Lacunarity                  1.918
#define     nOctaves                    8
#define     MultifractalOffset          0.95
#define     RidgedGain                  1.3
#endif

struct GenerationConstants
{
    float scale;
    float noiseSeaLevel;
    float spacing;
    float4 patchCenter;
};

struct OutputStruct
{
    float4 pos;
};

StructuredBuffer<GenerationConstants> terrainGenerationConstants;
RWStructuredBuffer<OutputStruct> patchOutput;

[numthreads(threadsPerGroup_X,threadsPerGroup_Y,1)]

//We lookup the the index into the flat array by using x + y * x_stride

void CSMain (uint3 id : SV_DispatchThreadID)
{
    GenerationConstants constants = terrainGenerationConstants[0];

    float2 sampleCoord = float2((id.x + constants .patchCenter.x),(id.y + constants .patchCenter.y));

    #ifdef HYBRID
    float noise = hybridMultifractal(NoiseFrequency*sampleCoord, OneMinusFractalIncrement, Lacunarity, nOctaves, MultifractalOffset);
    #elif defined RIDGID
    float noise = ridgedMultifractal(NoiseFrequency*sampleCoord, OneMinusFractalIncrement, Lacunarity, nOctaves, MultifractalOffset, RidgedGain);
    #endif

    float height = constants .scale*max(noise, constants .noiseSeaLevel);

    float4 output = float4(id.x*constants .spacing, height, id.y*constants .spacing, 1);

    int outBuffOffset = id.x + id.y * nVerticesPerSide;
    patchOutput[outBuffOffset].pos = output;
}

So once the ComputeShader completes, outputBuffer will contain the vertex position data. Not depicted here is normals generation, UV coordinates generation, color generation, etc.

Important - I do not try to draw a patch during the same frame in which it was generated. This has caused stalls for me in the past, and I don’t mind waiting a frame. Granted, this was awhile ago when I was using OpenGL (not with Unity), so it might not be a factor here, but I don’t mind waiting a frame to use the patch.

Now that the patch has been generated, on each subsequent frame, and for each and every patch, I do the following:

material.SetBuffer("patchData", patchGeneratedDataBuffer);
Graphics.DrawMesh(dummyMesh, transform.localToWorldMatrix, material, 0, null, 0, null, true, true);

By calling SetBuffer we’re basically hooking-up this patch’s vertex data (generated by the ComputeShader) as an input to the rendering shader. That data will be used to modify the dummyMesh, seen below.

Note that it’s not efficient to call material.SetBuffer like this, because it forces a new Batch to be created (Unity innerworkings). If Unity’s MaterialPropertyBlocks class supported a setBuffer method, that wouldn’t be a problem, but it currently doesn’t. So you’re stuck with 1 patch per batch, which just adds a bit of GPU driver overhead / state thrashing. It’s probably not a huge deal, but definitely not as efficient as it should be.

Inside the "ProceduralMeshVertSurf" shader (loaded and set earlier), we have the following:

Shader "ProceduralMeshVertSurf" {
    Properties {
	_Color ("Color", Color) = (1,1,1,1)
	_colorDeepWater ("Deep Water", Color) = (0.03, 0.16, 0.35, 1.0)
	_MainTex ("Albedo (RGB)", 2D) = "white" {}
	_Glossiness ("Smoothness", Range(0,1)) = 0.5
	_Metallic ("Metallic", Range(0,1)) = 0.0
    }
    SubShader 
    {
	Tags { "RenderType"="Opaque" }
	LOD 200
	
    CGPROGRAM

        #define nVerticesPerSide 224.0
	#define SHOW_GRIDLINES

	#include "UnityCG.cginc"

	// Physically based Standard lighting model, and enable shadows on all light types
	#pragma surface surf Standard fullforwardshadows
	#pragma vertex vert

        struct appdata_full_compute {
            float4 vertex : POSITION;
            float4 tangent : TANGENT;
            float3 normal : NORMAL;
            float4 texcoord : TEXCOORD0;
            float4 texcoord1 : TEXCOORD1;
            float4 texcoord2 : TEXCOORD2;
            float4 texcoord3 : TEXCOORD3;
        #if defined(SHADER_API_XBOX360)
            half4 texcoord4 : TEXCOORD4;
            half4 texcoord5 : TEXCOORD5;
        #endif
            fixed4 color : COLOR;
        #ifdef SHADER_API_D3D11
            uint id: SV_VertexID;
        #endif
        };

	#pragma target 5.0
        sampler2D _MainTex;

	#ifdef SHADER_API_D3D11
	  StructuredBuffer<float4> patchData;
        #endif

        struct Input {
	    float2 uv_MainTex;
	};

        void vert (inout appdata_full_compute v, out Input o) {
            #ifdef SHADER_API_D3D11

            float4 position = patchData[v.id];
            v.vertex = position;

            o.uv_MainTex = v.texcoord.xy;

            #endif
        }

	half _Glossiness;
	half _Metallic;
	fixed4 _Color;

	void surf (Input IN, inout SurfaceOutputStandard o) 
        {
	    #ifdef SHOW_GRIDLINES
	        float2 fract = fmod(IN.uv_MainTex*nVerticesPerSide, float2(1,1));
                fixed4 gridLine = any(step(float2(0.9,0.9), fract));
            #else
                fixed4 gridLine = 0;
            #endif

            #ifdef SHOW_PATCH_BORDER
                fixed4 patchBorder = IN.onBorder;
            #else
                fixed4 patchBorder = 0;
            #endif

            // Terrain color comes from a texture tinted by color
            fixed4 terrainColor = tex2D (_MainTex, IN.uv_MainTex) * _Color;

            fixed4 c = terrainColor + gridLine + patchBorder;

	    o.Albedo = clamp(c.rgb, fixed3(0,0,0), fixed3(1,1,1));

	    // Metallic and smoothness come from slider variables
	    o.Metallic = _Metallic;
	    o.Smoothness = _Glossiness;
	    o.Alpha = c.a;
	}
    ENDCG
    } 
    FallBack Off
}

So the real magic happens under void vert (inout appdata_full_compute v, out Input o) {...}, where the data stored in the patchGeneratedDataBuffer is used to modify the position of the dummyMesh’s vertices.

So there you have it! Or at least the basic process. Happy to answer questions as they come up.

1 Like

Glad to see I wasnt completely on the wrong track in whats going on. Though your strategy seems far more efficient and straightforward, and I wasn’t aware that you can do a Graphics.DrawMesh(…) operation with a dummy mesh.I am right now trying to reimplement this. Suffering from rendering the final plane yet, issue happens while I replace the vertices with

float4 position = patchData[v.id];
v.vertex = position;

in the vertex shader. If I remove these lines the prototype mesh is visible. But using the input buffer for the replacement does not yet work. Expect when I replace the vertice positions within the vertex shader using a noise function (and not the buffer). If I add the vertices from the buffer to the mesh’s vertices by v.vertex += position; , there is some distortion going on.

So I guess the vertex shader is OK but the issue is within reading from the buffer in the vertex shader, or probably more certain within the compute shader. However I am continuing finding the issue. Thanks for this deep insight in your latest post!

Edit: Finally found the issue. It was due to the fact that I didnt size the ComputeBuffer for the output right. I was used to think in Vector3s so I used 12 bytes for the buffer, but didnt realize that I was working on float4 so that I needed to use 16 bytes. But I think its been a good thing to spend 2 days for bugfix searching, so I was forced to to revisit everything and think about what I was doing. Well then, time for the next step, so look how and which shader to best rotate the plane / vertices to their required position/rotation so that I can create a cube with six of those (will review your posts, I am sure you’ve adressed this already).

1 Like

Well getting somewhere. Summary, what do we have right now?

  • A prototype mesh for later replacement in the vertex shader.
  • A constants ComputeBuffer that holds a number of information (nVerticesPerEdge, scale, spacing, patchCenter, planetRadius, noiseSeaLevel) that is being passed to a Compute Shader.
  • A compute shader that creates a plane Z-Up. nVerticesPerEdge * nVerticesPerEdge (so 224x224). Additionally I store all noise results as float in a separate additional buffer (I thought they still could be of use later).
  • A vertex shader that replaces each vertex of the prototype mesh with the
    Note: for each vertex being created in the compute shader I do the following:
    CPU:
    scale = 2.0f / nVertsPerEdge;
    GPU (ComputeShader):
    float4 output = float4(-1+id.xconstants.spacing, -1+id.yconstants.spacing, -1, 1);
    output.xyz = normalize(output.xyz);
    So that each plane is within [-1|1] range and normalizeable. Afterwards the noise is being added.

In order to proceed I have some very short design questions:

  • As right now each patch is created within [-1|1] range and from [-1,-1,-1] to [+1,+1,-1], where are they “shrinked” to take place within the right area (that gets smaller with each quadtree split)? Compute Shader or Vertex Shader?
    Or is the idea simply to decrease the spacing in the constants struct (eventually together with not start rendering from -1+id.x/y)?
  • Are the vertices “brought” to their final position (so basically rotated around (0,0,0) in order to form a quadsphere) within the Compute Shader or within the Vertex Shader?

Of course I could pass over the four edge informations of the quadtree nodes to the generation compute shader and directly place the vertices were they meant to be, but I guess this is not the original idea of instacing the planes and have them all Z-Up first. So it seems there now needs to be some shrinking and rotating being done next.

  • In order to calculate the normals my only idea right now is to have a second pass / second kernel program, “STAGE 2” that runs after all vertices were produced in “STAGE 1”. Which now can check the “surrounding” vertice positions from the buffer. Other less much effective idea is to do additional noise calls. Is there another approach?

  • The surface shader right now uses for testing purposes one texture, _MainTex. In order to have different textures (grass, stone, etc.), is it common style to create one main large atlas texture that holds all kind of terraintypes, or would one better create separate textures per type? Eg

    _GrassTex(“Albedo (RGB)”, 2D) = “white” {}
    _StoneTex(“Albedo (RGB)”, 2D) = “white” {}

2 Likes

Ugh, I can’t tell you how many times this kind of error (“Stride mismatch”) has got me!

Anyhow, congrats! Nice work! You’re quickly out-pacing me (time to quit my day job :slight_smile: )…

So before I address your questions, I need to apologize once again. You have ‘caught up’ to me so quickly that you’re now asking questions about concepts that are not yet incorporated into my “full” Unity project. That project has been been broken (non-compilable) for several weeks as I messed up a large refactoring in an attempt to break things down into multiple classes.

Thus, in order to give you appropriate answers to previous questions, I’ve been digging through an earlier Unity prototype which did not handle an entire planet, only one face of a cube. I’ve also been referring to some even older code from my first planet generator in OpenGL, which rendered an entire planet, but did things a little differently than we’ve been talking about (not large changes, but some subtle ones).

In piecing together the old code with some new concepts, I thought I could ‘smooth over’ the differences. But in reading-back, I’ve noticed several contradictions in my methods. This has probably caused some confusion. So, I must apologize for this in the event that it led you down the wrong path.

Thus In order to answer your most recent questions, I spent a couple of hours today “playing catch-up”, to incorporate the more-complete concepts from my older OpenGL planet and from my design notes into the “full” Unity Project. This project still does not compile, however, due to the botched refactoring I mentioned earlier, and so it remains untested. So the point is: there very well may be bugs in the code.

(I’m considering starting the Unity project from scratch, now that I’ve come up with a better class-design, it might be easier to build up from nothing than continue trying to refactor).

So that having been said, here’s my “actual” ComputeShader (only the relevant parts of the main function), which demonstrates patch generation within the scope of a full planet. I’ve filled in comments to explain the code line-by-line.

void CSMain (uint3 id : SV_DispatchThreadID)
{
      (...)

    // id.x and id.y each run from 0 to nVertsPerSide-1.  id.z == 1 (unused);

    // cubeFaceEastDirection is a vector indicating the cube face's "east" direction
    // cubeFaceNorthDirection is a vector indicating the cube face's "north" direction 
    // Each of these are either (1,0,0), (0,1,0), or (0,0,1), thus each vector is orthogonal to the other.

    // The cube is centered at (0,0,0) and the center of each face is a distance 'planetRadius' from the origin
    // Thus the cube will perfectly inscribe a flat sphere the size of the planet

    // patchCubeCenter is at the center of a quadtree node, sitting on the cube face.

    // 'spacing' is the distance between each vertex in a patch on the face of the cube
    // so for the largest patch (where 1 patch == 1 cube face), spacing === 2*planetRadius / nVertsPerSide
    // spacing is divided in half for each subsequent quadtree subdivision

    // terrainMaxHeight is max height from sea level (i.e. ~ 12 km)

    // first calculate the 'cube space' coordinates of the vertex:

    float eastValue=  id.x - (nVertsPerSide/2.0);
    // eastValue now ranges from -.5*nVertsPerSide to +.5*nVertsPerSide

    eastValue *= constants.spacing
    // eastValue now ranges from -.5*patchWidth to +.5*patchWidth (patchWidth is not actually a defined variable)

    float3 cubeCoordEast = constants.cubeFaceEastDirection * eastValue;

    // now do the same for the "north" direction:
    float northValue=  (id.y - (nVertsPerSide/2.0)) * constants.spacing;
    float3 cubeCoordNorth = constants.cubeFaceNorthDirection * northValue;

    // Now we'll generate the "cube space" vertex coordinate (i.e. restricted to surface of a planet-sized cube)
    // remember, cubeCoordEast and cubeCoordNorth are orthogonal to each other
    float3 cubeCoord = cubeCoordEast + cubeCoordNorth + constants.patchCubeCenter);

    // as an example, if:
    // cubeFaceEastDirection = (1,0,0)
    // cubeFaceNorthDirection = (0,1,0)
    // Then we're dealing with the cube face whose "normal" direction is (0,0,1);
    // Thus patchCenter will, by definition, be somewhere along ( ... , ... , planetRadius)
    // and so:
    // cubeCoord.x will range from [-.5*patchWidth + patchCenter.x, +.5*patchWidth + patchCubeCenter.x]
    // cubeCoord.y will range from [-.5*patchWidth + patchCenter.y, +.5*patchWidth + patchCubeCenter.y]
    // cubeCoord.z will equal planetRadius (for ALL vertices on this patch)

    // To reiterate, at this point 'cubeCoord' is a point restricted to the surface of a planet-sized cube.

    // now we determine the 'planet-space' value for the patchCubeCenter:
    float 3 patchCenter = normalize(constants.patchCubeCenter) * constants.planetRadius;
    // patchCenter now sits on the surface of a planet-sized sphere.

    // next, we normalize the patch
    float3 patchNormalizedCoord = cubeCoord.normalize();  

    // and then calculate its 'real world' size:
    float3 patchCoord =  constants.planetRadius * patchNormalizedCoord;

    // next we 're-center' the patch coordinates to the patchCenter - this is for our 'moving origin' (i.e. camera-centered)
    float3 patchCoordCentered = patchCoord - patchCenter;

    // next we generate the noise value using the patch's 'real-world' coordinate (patchCoord)
    // note: MultifractalOffset (a 3D vector) is used to offset the patch coordinates uniformly, so different planets can be generated.
    // this can be thought of as a different 'input seed' to the noise function.
    // also note: NoiseFrequency is probably << 1.0 (i.e. .0001). This reduces the size of the input coordinates to a size which
    // should not be affected by floating point precision issues
    float noise = hybridMultifractal(NoiseFrequency*patchCoord, OneMinusFractalIncrement, Lacunarity, nOctaves, MultifractalOffset);

    // the hybridMultifractal function (custom function) returns values from 0-1, so they must be scaled to the appropriate terrain height
    float terrainHeight = (noise * 2) - 1;  // terrainHeight now ranges from -1 to + 1;
    terrainHeight *= constants.terrainMaxHeight; // terrainHeight now ranges from -terrainMaxHeight to +terrainMaxHeight. 

    // this final step adds (or subtracts) the real terrain height from the real world-sized (but centered) patch.
    patchCoordCentered += patchNormalizedCoord*terrainHeight;

    int outBuffOffset = id.x + id.y * nVerticesPerSide;
    patchOutput[outBuffOffset].pos = float4(patchCoordCentered.x, patchCoordCentered.y, patchCoordCentered.z, 1);
}

So the end result is a terrain patch that, when translated to “patchCenter” will be located (and oriented) properly in ‘planet-space’ (i.e. planet centered at 0,0,0 world space).

This result is different than what I showed before, and so again, sorry for the confusion. I should have done this ‘translation’ into the new system first, before posting my earlier examples :flushed:

The first question is no longer valid, but the second question is: the vertices are ‘brought’ to their final rotation and scale within the Compute Shader - but the final translation (by patchCenter) must take place in the vertex shader. Again, that is so that the patch can be translated based upon a moving origin (i.e. camera-centered origin).

The vertex shader, then, becomes:

float3 patchCenter;

void vert (inout appdata_full_compute v, out Input o) {
            #ifdef SHADER_API_D3D11

            float4 position = patchData[v.id];
            // translate the patch to its 'planet-space' center:
            position.xyz += patchCenter;

            v.vertex = position;

            o.uv_MainTex = v.texcoord.xy;

            #endif
        }

Of course, note that I am not handling the “moving origin” (nor the location of the planet, for that matter), in the vertex shader. In my old OpenGL planet, I did so - but depending on how you set up Unity and your game objects, you can perform these translations entirely on the CPU. The surface shader applies all of the appropriate gameObject transformation “Behind the scenes” (somewhere in between the vert function and the surf function) according to the GameObject transforms. CAVEAT: to avoid precision issues, you’ll need to store camera, planet, and patch locations as doubles, then do the appropriate subtractions as doubles and finally down-sample to floats before passing the ‘final’ translation on the the GPU. And I believe Unity only stores the transforms as floats…so you may have to manage your own double-precision transformation on the CPU. I haven’t gotten this far with Unity, so I can’t give details.

But the reason that the patch is stored in the compute buffer ‘re-centered’ to the ‘patchCenter’ is to reduce the sizes of the floating point numbers. If you store the patch vertices as ‘real world’ positions with an origin of (0,0,0), (i.e. the vertices are planetRadius +/- maxTerrainHeight from the origin), then you’ll experience floating-point precision errors. Instead, by using the ‘patchCenter’ as the origin, then the vertices are only (at most) sqrt(2*(patchHalfWidth*patchHalfWidth) + terrainMaxHeight*terrainMaxHeight) from the origin.

In my OpenGL code, I performed “STAGE 2” as you describe, using transform feedback (a less-flexible type of Compute Shader) to numerically calculate the (approximate) derivatives (slope) of the terrain patch. It turns out you may not need to do this, however, because you can get the derivative of a simplex noise function analytically. See this: http://staffwww.itn.liu.se/~stegu/simplexnoise/DSOnoises.html …But, the math might get a bit complex, since you’d need to factor-in the curvature and orientation of the terrain patch. I have not yet done this, and don’t know if it would be more efficient than running a second pass (but my guess is that it would be more efficient).

I haven’t gotten that far with the Unity project - I’m also using just one Texture. (And BTW, my texture coordinates aren’t correct, either, because they don’t wrap properly at face edges - I’m considering using tri-planar mapping, but that approach is a bit expensive and can produce ‘blurriness’ when blending between textures. That’s on my to-do list. :confused: ) But to your question, atlases are more efficient in general, but I’m not sure how Unity’s back-end handles them. Plus if your textures are large you might exceed the GPU’s max texture size by packing them into an atlas.

With regard to selecting the appropriate texture, I believe the ‘material type’ should be stored on a per-vertex basis - at least, this is what I did in my OpenGL project. I determined the appropriate texture type in STAGE 2, since I had slope information in that stage, and used that information to select a different texture (i.e. rocky texture if slope is greater than a certain value).

Long post. Hope it’s helpful. Final apologies for earlier confusion. From here forward, I’ll refer only to my ‘current’ approach to avoid mixing concepts as I did earlier. But by the looks of it, soon enough you’ll be helping me!

Navy

Hi NavyFish,

thank you for the update. I implemented the new strategy ~2-3 weeks ago, unfortunately I did a “code-all-at-once-and-test-afterwards” approach and it didnt work :yum: . Well could’ve been anything, so today I started again building up on my working implementation, changing and checking everything step by step.

And yes, the extended constants buffer and compute shader, following your strategy, now works :grinning: . Very nice. Looking at it, it should be straightforward to append this approach to the six quadtrees for creating a sphere, as in the end it seems it is all about changing the parameters in the constantsbuffer per quadtreenode, the rest is done in the computer shader.

There is one issue left right now, as the generated mesh only reflects light when the light source is behind the plane, as seen on the screenshot. Not sure why this happens, but I havent yet implemented normals calculation, so we’ll see. But as soon as I fixed this, the plan is to append the plane creation in this shader to the six quadtree nodes next, to have a sphere I can work on.

2 Likes

Since you made it to shaders, here is a Logarithmic Depth Buffer implementation for Unity, made from Outerra blog post, might come in useful. There is the issue that Unity Surf Shader doesn’t accept a depth value as input, so the implementation is a Frag one instead of Surf.

Anyway…
http://pastebin.com/FJfEnWy7

Also, for testing normals, just normalize your vertex position if the planet is centered at zero, that will give you a perfect sphere’s normals.

That would give normals for a sphere, yes, but wouldn’t account for any terrain features.

Regarding lighting behind the plane, it sounds like the triangle winding order might be reversed. Hard to say for sure, not super familiar on Unity’s requirements.

Great work! I’m writing from the road so forgive the short response. Looking forward to tracking your continued progress!

Almost there. The lightningt now works after I implemented the normal calculation in the second pass / kernel in the compute shader. For the moment I did the simple approach (just use normalized coordinates for spherical lightning), as soon as the full sphere is generated I’ll be continuing in implementing the more accurate approach calculating the normals on their neightbours to take the noise into account.

I currently work on creating multiple planes (6, for a sphere). Right now I try to figure out if there is some dependency on when the buffers and materials are being created, the DrawMesh call done and if I need a Unity3D gameobject (as right now no two meshes are rendered at once) per DrawMesh call.
So right now my (not working) approach is I moved the buffers into the QuadtreeTerrain class (a quadtree node), as well as the material (not sure if individual materials are necessary).

class QuadtreeTerrain {
    // Quadtree classes
    public QuadtreeTerrain parentNode; // The parent quadtree node
    public QuadtreeTerrain childNode1; // A children quadtree node
    public QuadtreeTerrain childNode2; // A children quadtree node
    public QuadtreeTerrain childNode3; // A children quadtree node
    public QuadtreeTerrain childNode4; // A children quadtree node
    // Buffer
    public ComputeBuffer generationConstantsBuffer;
    public ComputeBuffer patchGeneratedDataBuffer;
    // Material
    public Material material;
        ....
}

In the SpaceObjectProceduralPlanet script, applied to a single game object, I hold six instances of quadtrees [=QuadtreeTerrain] then.

public class SpaceObjectProceduralPlanet : MonoBehaviour {
    ....
    // QuadtreeTerrain
    private QuadtreeTerrain quadtreeTerrain1;
    private QuadtreeTerrain quadtreeTerrain2;
    private QuadtreeTerrain quadtreeTerrain3;
    private QuadtreeTerrain quadtreeTerrain4;
    private QuadtreeTerrain quadtreeTerrain5;
    private QuadtreeTerrain quadtreeTerrain6;

    // We initialize the buffers and the material used to draw.
    void Start()
    {
        ...
        // QuadtreeTerrain
        this.quadtreeTerrain1 = new QuadtreeTerrain(0, edgeVector1, edgeVector2, edgeVector3, edgeVector4, quadtreeTerrainParameter1);
        this.quadtreeTerrain2 = new QuadtreeTerrain(0, edgeVector2, edgeVector5, edgeVector4, edgeVector7, quadtreeTerrainParameter2);
        this.quadtreeTerrain3 = new QuadtreeTerrain(0, edgeVector5, edgeVector6, edgeVector7, edgeVector8, quadtreeTerrainParameter3);
        this.quadtreeTerrain4 = new QuadtreeTerrain(0, edgeVector6, edgeVector1, edgeVector8, edgeVector3, quadtreeTerrainParameter4);
        this.quadtreeTerrain5 = new QuadtreeTerrain(0, edgeVector6, edgeVector5, edgeVector1, edgeVector2, quadtreeTerrainParameter5);
        this.quadtreeTerrain6 = new QuadtreeTerrain(0, edgeVector3, edgeVector4, edgeVector8, edgeVector7, quadtreeTerrainParameter6);
        CreateBuffers(this.quadtreeTerrain1);
        CreateBuffers(this.quadtreeTerrain2);
        CreateBuffers(this.quadtreeTerrain3);
        CreateBuffers(this.quadtreeTerrain4);
        CreateBuffers(this.quadtreeTerrain5);
        CreateBuffers(this.quadtreeTerrain6);
        CreateMaterial(this.quadtreeTerrain1);
        CreateMaterial(this.quadtreeTerrain2);
        CreateMaterial(this.quadtreeTerrain3);
        CreateMaterial(this.quadtreeTerrain4);
        CreateMaterial(this.quadtreeTerrain5);
        CreateMaterial(this.quadtreeTerrain6);
        Dispatch(this.quadtreeTerrain1);
        Dispatch(this.quadtreeTerrain2);
        Dispatch(this.quadtreeTerrain3);
        Dispatch(this.quadtreeTerrain4);
        Dispatch(this.quadtreeTerrain5);
        Dispatch(this.quadtreeTerrain6);
}

    // We compute the buffers.
    void CreateBuffers(QuadtreeTerrain quadtreeTerrain)
    {
        .... preparing generation constants
        quadtreeTerrain.generationConstantsBuffer.SetData(generationConstants);
        // Buffer Output
        quadtreeTerrain.patchGeneratedDataBuffer = new ComputeBuffer(nVerts, 16 + 12 + 4 + 12); 
    }

    //We create the material
    void CreateMaterial(QuadtreeTerrain quadtreeTerrain)
    {
        Material material = new Material(shader);
        material.SetTexture("_MainTex", this.texture);
        material.SetFloat("_Metallic", 0);
        material.SetFloat("_Glossiness", 0);
        quadtreeTerrain.material = material;
    }

    //We dispatch threads of our CSMain1 and CSMain2 kernels.
    void Dispatch(QuadtreeTerrain quadtreeTerrain)
    {
        // Set Buffers
        computeShader.SetBuffer(_kernel, "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
        computeShader.SetBuffer(_kernel, "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
        // Dispatch first kernel
        _kernel = computeShader.FindKernel("CSMain1");
           computeShader.Dispatch(_kernel, THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
        // Dispatch second kernel
        _kernel = computeShader.FindKernel("CSMain2");
        computeShader.Dispatch(_kernel, THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
    }

    // We set the material before drawing and call DrawMesh on OnRenderObject
    void OnRenderObject()
    {
        this.quadtreeTerrain1.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain1.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain1.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain2.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain2.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain2.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain3.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain3.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain3.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain4.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain4.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain4.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain5.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain5.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain5.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain6.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain6.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain6.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);
    }

    //When this GameObject is disabled we must release the buffers.
    private void OnDisable()
    {
        ReleaseBuffer();
    }

    //Release buffers and destroy the material when play has been stopped.
    void ReleaseBuffer()
    {
        // Destroy everything recursive in the quadtrees.
        this.quadtreeTerrain1.generationConstantsBuffer.Release();
        this.quadtreeTerrain1.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain2.generationConstantsBuffer.Release();
        this.quadtreeTerrain2.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain3.generationConstantsBuffer.Release();
        this.quadtreeTerrain3.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain4.generationConstantsBuffer.Release();
        this.quadtreeTerrain4.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain5.generationConstantsBuffer.Release();
        this.quadtreeTerrain5.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain6.generationConstantsBuffer.Release();
        this.quadtreeTerrain6.patchGeneratedDataBuffer.Release();
        DestroyImmediate(this.quadtreeTerrain1.material);
        DestroyImmediate(this.quadtreeTerrain2.material);
        DestroyImmediate(this.quadtreeTerrain3.material);
        DestroyImmediate(this.quadtreeTerrain4.material);
        DestroyImmediate(this.quadtreeTerrain5.material);
        DestroyImmediate(this.quadtreeTerrain6.material);
    }

    void Update() {
        // Do nothing
    }

}

Of course this is very bruteforce, but well this should work before I proceed as I need to figure out how to handle the buffers and draw calls and where to put them. I am a bit afraid I could need one gameobject per DrawMesh call, because I was hoping I could avoid multiple gameobjects (it may be that it could make sense anyway when I want to continue with colliders sight). Will investigate on that and edit this post then as soon as this is fixed. Then the fun stuff (refine the noise for a planet like visual and texture it) should come next, as well as scaling things up to 1:1 scale (right now I am on 1:1000 for debug purposes in the editor).

1 Like

Well let me also share my progress so far, I have been doing planet rendering since quite a while, I did my first poorly written planet rendering algorithm in 2011 and many versions later, but this is the first time I was able to achieve a game ready performance of a planet renderer. “Its getting late here, so i will detail my algorithm hopefully in the morning or sometime when i am free in coming days” But here are some of the screenshots of the current work in progress.

Planar Terrain works perfect, though spherical version of the algorithm is currently a work in progress as there are some bugs in it. In a simple summary my algorithm is a combination of quadtree and toroidally updating clipmap grids with respect to the level of detail. Basically i use quadtree to generate terrain patch origins which are then compared with the origins of patches in a list of clipmap grid patches and rendered if the origins match. And using the camera position I update the clip map grid patches toroidally which allows me to reuse the same memory without requiring any new memory creation and deletion at runtime.

Let me also share some of the previous works with my planar terrain algorithm which basically runs purely on CPU, it runs at 60 fps but there are frame rate jumps. Though the current algorithms shown above all run on gpu.








I also played with some of the unity’s default image effects such as color grading and tonemapping, to enhance the look of the screenshots.

8 Likes

Beautiful stuff, yuneeb! I look forward to seeing more of your approach.

I tried to get closer to the problem of rendering multiple objects with DrawMesh. It seems to depend which “Dispatch()” call I do first in Start() to dispatch each computebuffer to the computeshader before it is sent to the vertex buffer. It seems the first ComputeBuffer.Dispatch() call overrules all following ones, as only the results from the first call are drawn (in my case, only the mesh of QuadtreeTerrain1).
Edit: To be more precise: Both meshes ares drawn but they seem to share the same locations and probably buffer.
I noticed that as the rendered triangles doubled with each Graphics.DrawMesh added.
Although from my point of unserstanding I am using different buffers in the separate QuadTrees, thus I’d expect that all are calculated separately in the computeshader and drawn separately in the OnRenderObject() function. I seem to be something missing when it comes to work with compute shaders and multiple buffers. Are there some kind of restrictions, e.g. that you can only once dispatch to a computeshader on the GPU within a limited time or frames, or something else?

yuneeb90: as already mentioned I really like your prototype. You CPU implementation looks nice with the different effects combined. Looking forward to where your GPU implementation progresses, and to share noise settings for the terrain etc.

1 Like

you need to intialize materials separately for every dispatch call. You may need to do it through script like this.

// Initialize a new material for every node / patch.

GridMaterial = new Material(GridShader);

// Update for every node / patch.

GridRendererCS.SetBuffer (0, "VertexBuffer", VertexCB);
GridRendererCS.SetBuffer (0, "NormalBuffer", NormalCB);
GridRendererCS.SetBuffer (0, "TexcoordBuffer", TexcoordCB);
GridRendererCS.SetBuffer (0, "TangentBuffer", TangentCB);
GridRendererCS.SetBuffer (0, "BinormalBuffer", BinormalCB);

GridRendererCS.Dispatch (0, 1, 1, 1);

GridMaterial.SetPass (0);
GridMaterial.SetBuffer ("VertexBuffer", VertexCB);
GridMaterial.SetBuffer ("NormalBuffer", NormalCB);
GridMaterial.SetBuffer ("TexcoordBuffer", TexcoordCB);
GridMaterial.SetBuffer ("TangentBuffer", TangentCB);
GridMaterial.SetBuffer ("BinormalBuffer", BinormalCB);

which means that you will not assign a new material through project view in unity rather you have to create a new material through script and to set material properties you will also need to do it through script for every material.

Thanky yuneeb90, appreciate the quick help,

EDIT [2015-11-25 08:33]: Guess I found it, as you made me try something out while looking at setting the buffer. Looks like you need to set the buffers again when you switch the kernel (sight :expressionless: ). After changing the Dispatch function, adding two lines to set the same buffer again after the first dispatch, it worked.

Seams betweens the patches appeared, need to look at the shader code again where the vertex positions are defined, but well, an important step forward.

//We then dispatch threads of our CSMain1 and CSMain2 kernel.
void Dispatch(QuadtreeTerrain quadtreeTerrain)
{
   // Set Buffers
   computeShader.SetBuffer(_kernel, "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
   computeShader.SetBuffer(_kernel, "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
   // Dispatch first kernel
   _kernel = computeShader.FindKernel("CSMain1");
   computeShader.Dispatch(_kernel, THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
   // Set Buffers
   computeShader.SetBuffer(_kernel, "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
   computeShader.SetBuffer(_kernel, "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
   // Dispatch second kernel
   _kernel = computeShader.FindKernel("CSMain2");
}

2 Likes

Writing from my phone and without access to the code, so forgive a short response. yuneeb is correct in that you need a separate material for each patch (for now, although I’ve submitted a feature request to Unity to allow MaterialPropertyBlock.SetBuffer() , which would allow you to reuse a single material and thus more efficient batching).

You should only need to search for the Kernels once, then save the results as a variable. Searching every time will have a significant performance impact.

And, just double checking: you should only have one ComputeShader object for each pass (and thus one kernel for each). Each patch would share these objects. Each patch needs its own material and own set of Buffers, of course.

Hi NavyFish, yuneeb,
followed your all your hints now (separate materials per patch, one compute shader object, moved the kernel search into the awake routine). The six patches now render quick and without problems.
EDIT 2:
Fixed the seam/space left between the patches.

I had to change the c# CPU code to

parameter.scale = 2.0f / (nVertsPerEdge);
parameter.spacing = 2.0f / (nVertsPerEdge-1.0f);

and the compute shader code to:

// First calculate the 'cube space' coordinates of the vertex:
float eastValue=  id.x - ((constants.nVertsPerEdge-1)/2.0); 
eastValue *= constants.spacing;                            
float3 cubeCoordEast = constants.cubeFaceEastDirection * eastValue;

// Do the same for the "north" direction:
float northValue=  id.y - ((constants.nVertsPerEdge-1)/2.0);
northValue *= constants.spacing;
float3 cubeCoordNorth = constants.cubeFaceNorthDirection * northValue;

I investigates further why my original problem occoured that not all planes were rendered, and found out it concentrated on the second kernel dispatch call. When I left the second dispatch call in, altough I did nothing yet in that second compute shader kernel, the patches did not render completely, especially the last one to be dispatched. When commenting the second kernel dispatch call out (which I do not need now but later for the normals), everything works withough any problems.

// Set Buffers
computeShader.SetBuffer(_kernel[0], "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
computeShader.SetBuffer(_kernel[0], "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
// Dispatch first kernel
computeShader.Dispatch(_kernel[0], THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
// Set Buffers
//computeShader.SetBuffer(_kernel[1], "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
//computeShader.SetBuffer(_kernel[1], "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
// Dispatch second kernel
//computeShader.Dispatch(_kernel[1], THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);

Makes me come to an important question:
Does the CPU stall and wait while the GPU works in the compute shader for the first dispatch? Or does it work async and the CPU immediately dispatches to the second kernel, no matter if the first kernel is done yet?
If the second is true, then I would have underestimated the orchestration of dispatches in the CPU and it would be clear why the above code wouldnt work.

Slowly catching up, fBm (Fractal Brownian Motion) added and coloring based on noise results. Right now I decide the coloring and so the terrain type completely in the surface shader. Will move this into the compute shader (decision of terrain type) once I figured out how to do the second compute shader kernel dispatch correct. Not anywhere as beautiful as your implementations, NavyFish and yuneeb90, but anyway happy to getting used to shaders more and more.

Of course I have yet no good indication of the performance, gain, however I am already impressed how fast the terrain renders even at very high octaves. Wow :open_mouth: . Looking forward to reimplement the quadtree LOD soon (should be more or less a copy-paste takeover of my old code fingers-crossed) and see how the performance behaves and continue playing with noise and increased terrain detail.

Going to extract the code to create a small Unity3D project for upload/link here now.
Would have liked to check how to do teh second dispatch to the compute kernel right, but anyway, besides that its pretty complete now to cover the whole topic initially, the small project should be a good starter fo the next one who needs to go the same route like me.

1 Like

@JoergZdarsky use an offset while generating vertices. I did it like this

x and z are indices in a single patch.

vx = (-HalfLodSize + EdgeLength * (x + 0.5f));
vz = (HalfLodSize - EdgeLength * (z + 0.5f));

this will perfectly align your grid patches. you will though need to see according to your algorithm how you will add that offset.

1 Like

The second is true.

All work on the GPU is done asynchronously from the CPU, but any command you send to the GPU is executed in the same sequence that you sent it. Sometimes half of the GPU will be working on one command and the other half on the next command. When the next command absolutely depends on the previous command having been completed, the graphics driver will ‘flush’ the GPU’s command queue up to that point, forcing it to catch up. This can manifest in GPU stalls.

But the CPU will rarely stall waiting for the GPU. This only happens when a resource must be synchronized between the two, such as during a CPU ‘read back’ of GPU data. In that case the GPU will flush, and the CPU will ‘block’, waiting for the GPU to finish and grant access to its resource.

Anyway, you should ‘ping pong’ or ‘double buffer’ your 2nd stage Compute Buffer data. In other words, don’t overwrite data in a buffer that was generated during the same frame. This will force the GPU to finish the first task before beginning the second task, thus destroying parallelism.

Another thing to consider is that switching between Kernels or Compute Shaders is expensive - so do all of your stage 1 processing first, and then switch to stage 2 and do all of that work, instead of swapping back and forth between stage 1 and stage 2 for each patch.

Another thing to look at in the Compute Shader’s code is StructuredBuffer vs RWStructuredBuffer. The former is read-only.