Procedural Terrain Rendering How-To

The thing to understand about the GPU is that it’s essentially a massively multi-core machine. Like 2048 cores (on the latest GTX 980), or 1024 cores on a more modest NVidia card. Each one of those cores can execute a fairly complex instruction set, although not quite as complex as a modern CPU.

These cores must all share memory, which comes in different forms, ranging from “Global” to “thread group”. Without getting into all of those details (not particularly important at this moment), the point is that even at the most granular level of memory, it is shared by ~48 cores. Thus, your code must be written in a parallel manner, which is foreign to most developers used to writing code on the CPU. Check out some tutorials on “Prefix Scan” or “Parallel Prefix Sum” to give you a sense. But essentially this boils down to re-writing your algorithms so that chunks can be worked on simultaneously. The simple example in a serial context: for(int i=0, i<100, i++){ sum += data[i]; } This will clearly sum up an array. But it uses a single thread. To do this in parallel (and assuming you have at least 50 threads):

first iteration:
for each thread i, where i >= 0  and i < 50
output1[i] = data[i*2] + data[i*2+1]

second iteration:
for each thread i, where i >= 0  and i < 25:
output2[i] = output1[i*2] + output1[i*2+1]

third iteration:
for each thread i, where i >= 0  and i < 13:
output3[i] = output2[i*2] + output2[i*2+1]

et cetera, until after 7 iterations you’re left with a single piece of data, which is your answer. (Note you don’t actually need 7 output buffers… you can use simple ‘double buffering’ and just flip back and forth between them).

Actually, for something this simple, it would likely be much more efficient to do so using fewer iterations (and thus less threads overall, but more operations per thread) because there is a lot of overhead involved with setting up each iteration. Visually, here’s a good example of doing a sum with more operations per thread:

The 64 teal dots in this case represent your original data. Say each of them represent a single floating point number from -1 to +1. Lets put them in an array called data[]. Sum them as such:

first iteration:
spawn 16 threads, each labeled 'i', where i >= 0 and i < 16  <---- these are the red dots
output1[i] = data[i*4] + data[i*4+1] + data[i*4+2] + data[i*4+3]

second iteration:
spawn 4 threads, each labeled 'i', where i >= 0 and i < 13   <---- these are the blue dots
output2[i] = output1[i*4] + output1[i*4+1] + output1[i*4+2] + output1[i*4+3]

third iteration:
spawn 1 thread   <---- this is the topmost black dot
sum = output2[0] + output2[1] + output2[2] + output2[3]

Looks familiar, doesn’t it! Luckily for us graphics cards are very well suited to performing computations on large arrays (or n-dimensional matrices) of data, and are logically well-suited to spatial subdivisions like the quadtree above. Again, the key concept here is thread-independence. If your task can be broken into many subtasks which each run parallel and no not depend upon the (in-progress) results of each other, then chances are it can be well-adapted to the GPU.

Interruption to regularly scheduled programming: My wife is begging to play “Lovers in a Dangerous Spacetime” (hell yes my wife likes to play games… I’m the luckiest man on Earth). I’m going to finish this post here and pick it up tomorrow (day off! woo!). Sorry :slight_smile: But here’s what’s in the queue:

First I’ll discuss two aspects of GPGPU planet generation: terrain heightmap generation, and vertex normal generation. Heightmap generation is very straightforward, particularly because each vertex is independent from one another, and only needs one piece of information as input: position. Vertex normal generation is slightly more complex, however, because each vertex must read from multiple locations (i.e. its surrounding vertices). It’s not so bad, though, because this is a read-only operation, and thus the normal for every vertex can be calculated in a single pass.

Next I’ll start getting into what this looks like in Unity, with some code examples. I’ll probably not have fixed my main Unity project by that point, but do have a simpler terrain generation project which will be well-suited to the task (It’s basically a single behemoth monobehaviour, but I’ll break it down).

Till next time.

I think you’re absolutely right, how did I miss that. Guess my strategy (well, better “simple implementation”) works with normal spheres but not after they have been distorted by noise. Thus the current lightning is very subtile e.g. for the asteroid, and does not take the deformations into account and everything looks a bit washed out. Thanks for the hint on this one. This should be an important change, will see if I can find a generic implementation for the normals.

Besides that, THANK YOU for these really great latest posts. Impressive insight and lots of very interesting and valuable important information. Will take my time to take a deeper look into them, and start preparing myself by maybe create some more simple shaders to get used to them before doing the more complicated things. This thread is really one of the most valuable on the overall topic.

Absolutely, looking forward to get an inside of how you use Unity for these GPGPU calulcations in the shader!

Joerg, it’s been my pleasure to discuss these things ,and thank you for the patience. I’ve been really busy lately. Everything is culminating towards next September, when I’m finally done with my present employment and can take all the time I want to work on these hobbies. But hopefully I’ll post a follow-up before that time :smile:

For an algorithm to calculate the normals, refer to post #73 on this thread. If you’d like further information on the technique, share the relevant snippets of your code and I’d be happy to help you find the normals. But no doubt you can figure it out.

Shaders ARE modern graphics programming. Even the old ‘fixed function’ pipeline that was the norm for so many years is now just emulated using the programmable pipeline. So coming to understand those is quite important in my opinion. Perhaps in my next post I’ll give a description on what that actually means.

Unity certainly uses Shaders deep inside the engine (many very complex ones), and to some degree allows you to write your own code. But it’s very important to recognize that Unity provides a large middle layer between you the programmer and the actual Shader code that’s fed to the GPU. In fact, for a very simple Shader that you manually wrote, Unity recompiles tens of different iterations of that code, into both HLSL, GLSL (two shading languages), based upon different target hardware capabilities and developer-selscted functionality (ie is deferred rendering enabled? Are there shadows? Etc). No only that, but it does tons of ‘plumbing’ in the background to connect your Shader code to your host code (I.e. monobehaviours etc) that you never see.

There’s nothing wrong with that - in fact, its probably the key selling point of the language, and is certainly critical for cross-platform development. But I feel it’s important to understand the distinction… And personally my learning style is to understand the fundamentals first before diving into the more complex/abstracted implementation represented by Unity.

Anyway, Unity offers plenty of tutorials regarding how to wrote shaders, and I encourage you to explore those. They really do make things simple for 90% of the use cases… And thus implicitly make things impossibly difficult for the remaining 10% :slight_smile: (and GPGPU is definitely one of those edge cases).

Writing from my phone waiting for a meeting to start. Apologies for any typos. Till next time :slight_smile:

@NavyFish:
I started a new implementation of the procedural planet’s planes, as most likely I need to create a lot of new code when using shader, and to have something for practicing. For the shader tutorials I now render one flat plane to which I add custom shader. Trying to get used to shaders to hopefully be prepared when you give some more insights with your example implementation, and I dont need to look at it from scratch. I am now at the point where I create a plane out of 224x224 shared vertices, apply a custom shader, provide some values from Unity3D to the shader (e.g. 224 float width value, vector coordinates of the plane’s four edges) as I maybe need them later, and then have the shader paint everything of the plane in red. And read back one of the shaders custom textures (e.g. _heightmap) back to Unity3D at some point.

At the same time I try to switch to some of your techniques (e.g. instancing planes in the model space). Two questions rised here:

Instancing new planes:
You said you create the planes all the same way in the same coordinates [model space] and then rotate/translate them to the desired location of the planet afterwards. The planes normal would be Y-Up. So I implemented it with each plane instanciated with the edge coordinates [-1,0,+1], [+1,0,+1], [-1,0,-1], [+1,0,-1] (should make AABB tests easier). Is that layout right? Because two of your posts read different:

Number of vertices / shared vs unique vertices
While implementing this approach for e.g. 224x224 width = 50176 vertices to stay below the 65K limit, I realized that this means you need to share vertices for the triangles. In my latest implementation I used unique vertices for each triangle, which made the implementation more complicated and much more expensive of course, but I though you need to go this way for more lightning flexibility. I go the shared vertices route now, but I wonder if there wont be a limititation for sharp edges etc. doing it this way?

Nice, I think you’re doing the right thing learning to work with Unity’s shaders. To answer your questions:

1st question: If you go with y as “up”, then, you should do: [-1,1,+1], [+1,1,+1], [-1,1,-1], [+1,1,-1], and give every vertex in the plane a y-value of 1. This is because the actual scaling of the ‘prototype’ patch to the appropriately sized patch (based upon LOD) takes place in the shader, and thus the y=1 will be multiplied by R, and the spacing between each vertex (currently 2/224) will be multiplied by the patch’s ‘world space’ diameter.

I’ve actually come to prefer z being “up”, so my ‘prototype’ patch is an X-Y plane, where Z = 1, but either approach is just fine.

2nd question: Yes, you will have ‘smooth’ lighting when using shared vertices because there’s only 1 normal vector at that vertex (the average normal of the two adjacent triangles) instead of two separate normals. But in practice this is unnoticeable. It matters for simple geometry like a cube represented by 8 vertices. If you represent that cube so that each face has 224*224 vertices, you’ll have a very sharp line for lighting at the edges. And we’re not even yet talking about normal mapping.

Was just rereading an earlier post (post #101- the one about parallel processing algorithms) and noticed a significant errror in my pseudo code. Fixed, in case anyone refers to those snippets trying to learn about this stuff. I’m getting around to posting more details on the GPGPU process as promised, but life (well, job) has a way of keeping us from doing the things we love! (Until you either make job doing what you love (ideal), or make job go bye-bye (temporary solution))

Well now the Y-Up being [x,1,z] for all vertices makes sense now that you mention the re-scaling done on the shader. But now as you mention this, it makes me worry how much has to be done in fact in the shader and not on the CPU / in Unity c#. If I got this right and refering also the concept of model space, I consider the following tasks are to be performed which I right now have in place all on the cpu / in c#

CPU:

  • Create each prototype patch between [-1,1,+1], [+1,1,+1], [-1,1,-1], [+1,1,-1]. Each. All will remain there for the rest of the time.

GPGPU / Shader:

  • Rotate each patch (or each vertice?) to the appropriate position (hand over the center position from Quadtree to the shader and calulcate and use elevation and azimuth).
  • Normalize each vertice in the shader.
  • Rescale each patch to be within the boundaries of the four quadtree edges (hand over the quadtree edges or scalefactor or worldspacediameter) in the shader. So, multiply by R and change spacing between each vertex.
  • Calculate the noise (with a noise function within the shader code) for each vertice in the shader
  • Set the color for each vertice depending on the noise value.
  • Reposition each pixel depending on the noise value.

If this is the case, than I am already done with the cpu code hurray :wink: , but I need loooooots of shader howto with a required strong learning curve shock. Is that really how you devide the tasks?

If yes, than I should structure my first shader learning sessions into:

  • How to color a vertice depending on some value (e.g. each vertice position)
  • How to reposition the plane, e.g. rotate it at certain amount or move it to some different position.

Joerg,

You are a great student and I a horrible teacher. I apologize for being so absent from this thread as of late.

I agree, learn about how to manipulate vertices in the shader, to include color and repositioning as you mentioned. The thing to realize is that a vertex shader is run once for every vertex fed to it (and many in parallel). So instead of a shader having “loops” over all of the vertices in a patch, it acts as if there is only 1 vertex to manipulate. I’m actually doing patch heightfield generation in a Compute Shader, not a vertex shader, but for all intents and purposes (in this use case) they’re the same thing - it’s run once per each vertex in whatever geometry is fed to it (the input patch, for us).

You’re more or less right about the breakdown of duties between CPU and GPU, but there are a few things to note:

First off, on the CPU, you’re actually only ever creating a SINGLE patch (of 224x224 vertices). Not one per terrain patch, but one. This is the only thing stored in host RAM! (until much much later, caching data read-back from the GPU, but this is currently outside the scope and not something I have implemented). And secondly, yes, you are correct in sensing that most of the work is done on the GPU. But fear not - most of that work is only ever done once per patch.

You actually have two shaders: your patch generation shader, and your patch rendering shader. I actually break my generation stage into two phases, so use two separate generation shaders: one which builds the height field, and another which calculates the numerical gradient (and thus normals). But it turns out you can get the gradient of simplex noise at a point analytically fairly simply, so the process of generating normals can be tied in with the heigthmap generation into a single shader. I’m still going to describe the process by using 3 total shaders, though, because A) that’s what’s in my code at the moment, and B) this will demonstrate the process of using multiple shaders, which you may need to do later anyway (more complex procedural terrain algorithms probably have 3 or 4 stages - i.e. they might run a plate tectonic simulation for a few frames before generating the heightfield),

So the process goes something like:

CPU: (once)

  • Create a single prototype patch (mesh). Doing this in unity automatically “uploads” those vertex positions to the graphics card. This mesh is used when rendering - it serves as a blank ‘template’ geometry with properly configured triangle windings, etc. The vertex shader (GPU) during the rendering phase will modify the positions of the vertices in this mesh (every frame, but this is a very quick process), according to the heightfield data generated earlier by the GPU.

CPU: (every frame)

  • Determines (every time camera moves a certain distance) which patches need to be drawn.

  • Maintains cache of all currently generated patches (not the patch data itself, just the ‘name’ (coordinates) of the patch and other details). All patches currently being drawn are contained in this cache, but not necessarily all patches contained in the cache are drawn. The cache is a ‘Least Recently Used’ cache, so when a patch goes out of sight (i.e. too far away that the LOD switches), and thus is no longer rendered, it is still retained in the cache until either drawn once again or replaced to make room for new patches.

  • Checks the list of which patches need to be drawn, compares it against the list of those that are drawn (or are in the cache but presently not drawn), queues a list of patches to be generated and patches to be rendered.

  • Commands the GPU to generate requisite patches. This can be limited to preserve framerate (i.e. don’t try to gen more than 10 patches/frame).

  • Submits the single mesh for rendering, multiple times (once for each patch to be drawn), along with the associated GPU data buffer ‘handle’ for that patch (a pointer to GPU memory which stores the terrain’s height data, generated by the GPU - will explain more on this in a second).

GPU:

  • When commanded to generate a patch, it receives a small data buffer (a struct) from the CPU which describes the location/size/ etc of the patch. From there, the GPU executes a highly parallel piece of code in a Compute Shader, which runs the simplex noise calculations and generates a height value. The CPU tells this Compute Shader to execute 224 x 224 times (i.e. once for each vertex that needs to be generated). Once calculated, the height value for each vertex is stored on the GPU in a data buffer, which is just a contiguous block of VRAM. The CPU will get a ‘handle’ (pointer) to this block of VRAM (so it can refer to it), but does not have direct access to the data in the block. This only happens once per patch

  • If using multiple generation stages, once all patches have run through stage 1, the GPU loads up a new Compute Shader (kernel), and executes that one. The CPU has to ‘orchestrate’ this behavior, basically by setting up references: i.e. ‘Stage 2, get your input data from these specific buffers that were generated by Stage 1, then output your results to these buffers over here’. I will show code for that in another post, but it’s pretty straightforward. Switching shaders incurs a cost, however, so it’s sometimes useful to only perform a single stage per frame.

  • When it comes time to render a terrain patch, the CPU actually tells the GPU to render the prototype mesh geometry (which is completely flat). But it also points the GPU to the associated VRAM data buffer (pointer) for the correct patch that was filled earlier in the Compute Shader. During the vertex shader stage of the rendering process, the mesh’s vertex positions (and any other property, i.e. color) is modified by copying the values from the terrain patch’s data buffer the CPU pointed to. This is extremely quick, because that data is already stored on the GPU (data transport CPU <–> GPU is slow, but on-GPU memory access is quite fast), so it’s really just a memory read and then write. Thereafter, the shader goes about it’s normal routine of taking the vertex positions and applying matrix transformations to place them on screen where expected (due to camera position, etc). This ‘model-view-projection transformation’ already happens for every piece of geometry that is submitted to the GPU by Unity.

So the rendering process on the shader is trivial - it’s basically the same cost as rendering any other piece of geometry. It’s the generation stage that is somewhat expensive, but this only needs to be done once. And doing it on the GPU is Significantly faster than doing it on the CPU.

Generation details:

The CPU passes to the GPU the following information (as a struct) for each patch to be generated (though you can separate some of these into “Global” shader values if desired):
nVertsPerSide = 224 vertexSpacing = .8928 miles planetRadius = 4000 miles terrainMaxHeight = 8 miles patchWorldSpaceCenter (2d vector)(note for now I’m just assuming sea level is at “0” and the terrain will range from -8 to +8 miles)

First off, I calculate the patch’s world-space width (pre-normalization)
patchWidth == vertexSpacing*nVertsPerSide patchWidth = .8928*224 patchWidth = 200you could pass this in as a pre-calculated variable to be a bit more efficient, but the shader still needs access to vertexSpacing and nVertsPerSide anyway, and it only costs 1 multiplication

  • Next I normalize the patch - e.g. to turn a flat plane into a curved surface. I do this now because this step suffers from precision loss (when done in single-float precision, as double aren’t natively supported on most GPUs) which is generally acceptable and unnoticeable - but if done after the terrain heightfield generation, then the loss of precision is much more visible in the terrain features. Also at the end of this step, the input patch will be transformed into the correct 'world space ’ size.

    In this step, I refer to x_thread_index and z_thread_index. The important thing to know for now is that the Compute Shader is executed in a “block” of threads - in this case 224 x 1 x 224 threads (x, y, z, respectively). Each shader invocation has access to its corresponding index.
    Also, note that I’m going to refer to a variable ‘outVert’. This variable will eventually become the final “output” of the generation phase, which will be dumped into the memory buffer on the GPU, thus containing the final heightfield data (and anything else you’d like it to contain, i.e. color, normals, etc etc).
    float hw = patchWidth * .5; float patchXpos = ((x_thread_index / nVertsPerSide) * 2.0) -1; float patchZpos = ((z_thread_index / nVertsPerSide) * 2.0) -1; outVert.x = patchXpos * hw; outVert.y = planetRadius; outVert.z = patchZpos * hw; outVert.normalize(); outVert.y -= 1; outVert.multiply(planetRadius);
    Keep in mind that this program (the patch generation shader) runs once for each vertex in the input patch - so we don’t perform a ‘loop’ in the code. The shader just deals with a single vertex, and 224224 of these mini programs are run simultaneously (well, not unless unless you have a beastly GPU. more like 224224/#ofCUDAcores your gpu has).
    So this has the effect of curving the patch accordingly, and scaling it to the appropriate world-space size. The center of the patch (pre-noise) will be at 0,0,0.

  • Next, I determine the terrain generator’s input coordinates based upon the vertex patch’s X and Z input coordinates, the patch’s world-space center we passed in as well as the vertexSpacing:
    generatorInput.x = (outVert.x * vertexSpacing) + patchWorldSpaceCenter.x generatorInput.y = (outVert.z * vertexSpacing) + patchWorldSpaceCenter.z

  • Note, the patchWorldSpaceCenter is in pre-normalization “cube space”, ie the world-space coordinates of the ‘planet cube’ before it becomes normalized to fit a sphere. This is the center of a quadtree node.

  • Also, note that generator input is only 2-dimensional, so its inputs are technically “x” and “y”, even though the patch dimensions being fed into it are “x” and “z” (since Y is “up” for the patch)

  • These inputs are fed into the terrain generator (noise function), and stored to the variable height. But the output is in the range of +1 to -1, so we must scale it appropriately by multiplying it by terrainMaxHeight.
    ``

  • Finally, we apply this height value to the output vertex by scaling the vertex:
    outVert.multiply(height);
    Note that I’m not simply adjusting the vertex’s y-value, as this would be a mistake: we’ve already normalized our patch, so it’s a curved surface, and thus doing so would throw off the curvature. By multiplying, we’re effectively scaling the vertex along it’s own direction.

Okay!!! I’m going to stop here, because I’m tired, and because this is rapidly becoming a mega-post. Will pick up where I left off next time. Questions? Fire away…

1 Like

edit: whoa! lots of typos and a few big mistakes in the post, sorry. If you read it within 20 minutes of being posted, you may want to skim back over. Sorry!

edit#2: as of this edit, I changed the normalization step to be a little more clear (added descriptions for x_thread_index, etc. If you referred to the post prior to this edit, it’s worth a re-read of that section (look at the edit history so you only have to read the changes

1 Like

No need to apologize, your posts are full of valuable information and every single post is very much appreciated.
I am getting somewhere with some vertex/fragment-shader trial and error tests. Patch applied with a shader that adds simplexnoise based on the vertex positions (noise input).
Extending this to Fractal Brownian Noise with different octaves should be straightforward but isnt crucial yet.

Brute force coloring (not nice, just for testing). Combining this with different textures in the shader should also
be straightforward as the process in reading from textures seems simple.

Displacement of vertices based on the noise. Solved inefficient because I do one Noise call in the vertex shader and another call in the fragment shader. Anyway doesnt matter as this just for testing.

This is right now done with one Shader, called “Custom/ProceduralPlane” applied to this single patch. It has one SubShader were in the vertex section I do the displacement to the position of the vertex (1st noise call) and a fragment section (2nd noise call) that colors the vertex. Suprisingly simple until now. But…

[quote=“NavyFish, post:110, topic:765”]
First off, on the CPU, you’re actually only ever creating a SINGLE patch (of 224x224 vertices). Not one per terrain patch, but one. This is the only thing stored in host RAM!
[/quote].

This is the crucial part for me. Will stick to this in the next post.
Got used with Compute Shaders meanwhile and created the first plane with it (screenshot shows dots), with noise and displacement in the Compute Shader (input is a verticeBuffer out of the prototype plane) before handing the three buffers (vertice positions, triangles and noise) to the vertex shader for rendering.

EDIT: Post shortened, tldr :wink:

Hi NavyFish,

Can you please detail these sections?
What I right now do is:

  • Create a ComputeBuffer “verticePosBuffer” with 224x224 vertice-positions (input for this is the flat prototype plane)

  • Create a ComputeBuffer “trianglesBuffer” with the triangle indices (input for this is the flat prototype plane)

  • Create a ComputeBuffer “noiseBuffer” in which 224x224 float values are written into in the Compute Shader

    void CreateBuffers(){
    int VERTICESCOUNT = this.plane.vertices_pos.Length;
    int TRIANGLESCOUNT = this.plane.triangles.Length;
    verticesPosBuffer = new ComputeBuffer(VERTICESCOUNT, 12); // (Vector3 -> 12 bytes)
    verticesPosBuffer.SetData(this.plane.vertices_pos);
    trianglesBuffer = new ComputeBuffer(TRIANGLESCOUNT, sizeof(int));
    trianglesBuffer.SetData(this.plane.triangles);
    noiseBuffer = new ComputeBuffer(VERTICESCOUNT, sizeof(float));
    outputBuffer = new ComputeBuffer(VERTICESCOUNT, 12); //Vector3 -> 12 bytes)
    }

  • Pass all three buffers to a Compute Shader that
    – fills the “noiseBuffer” with a heightmap (224x224 floats)
    – displaces the vertices from the verticesPosBuffer and fills the outputBuffer with the displaced vertices.

    void Dispatch() {
    computeShader.SetBuffer(_kernel, “vertPosBuff”, verticesPosBuffer);
    computeShader.SetBuffer(_kernel, “noiseBuff”, noiseBuffer);
    computeShader.SetBuffer(_kernel, “output”, outputBuffer);
    computeShader.Dispatch(_kernel, 32, 32, 1);
    }

When this is done I pass the buffers to the render shader:

void OnRenderObject()  {
   int VERTICESCOUNT = plane.vertices_pos.Length;
   int TRIANGLESCOUNT = plane.triangles.Length;
   Dispatch();
   material.SetPass(0);
   material.SetBuffer("buf_vertices_pos", outputBuffer);
   material.SetBuffer("buf_triangles", trianglesBuffer);
   material.SetBuffer("buf_noise", noiseBuffer);
   Graphics.DrawProcedural(MeshTopology.Triangles, TRIANGLESCOUNT);
}

I understood that you don’t pass the vertice-positions to the compute shader but only the boundary information (width of plane etc.) to create the noiseBuffer. Makes sense, a switch to that should be straightforward. But I have a problem to understand how then proceed in rendering multiple planes out of one initial prototype plane from the CPU and how you pass positions to a vertex shader.

Do you ever create a verticesPositionBuffer like me and pass it over (once or many times for each plane) to a vertex shader?
Which information (buffers) do you pass to the vertex shader?
You said you create a mesh as workaround. Do you create for each plane a UnityEngine.GameObject or UnityEngine.Mesh to apply a stock surface shader? And if so, how do you displace the vertices with a custom vertex shader while you want to use a surface shader the same time?
I think seeing how you’d do the section “OnRenderObject()” would be very interesting.

Thanks a lot beforehand!

PS: Pretty excited to get plane creation on the GPU running and then combine this with the current implementation. This parallel processing of vertices (I did 32x32 planes before in serial on the CPU) is really impressive and along with the shared vertices and leaving everything on the GPU should give a massive performance boost. Also, when switching to 224x224 planes I might skip a few highest levels of my quadtree as there is already enough precision on lower nodes.

Can’t wait to give a detailed response. I’ve been out of town all weekend and am traveling home today.

Great job figuring out how to use Compute Buffers in Unity! I’ll have to refer to my code to address a few details but you’re definitely on the right path.

You’re correct that I don’t pass a base plane or patch to the Compute Shader’s - just basic information relevant to that patch’s generation, such as size and position on the planet. And correct regarding the use of Compute Buffers, but the only ones I fill are position, normals, and (eventually, but not in my current implementation) texture coordinates. Position represents the final vertex positions, ie scaled appropriately for the LOD, curved to match the surface of a sphere, and petrurbed by a noise function.

The mesh (not procedural mesh) I use already has its triangles array defined. You only ever need to do this once. I draw the mesh normally, but give it a custom surface shader, which has a custom ‘per vertex pass’ (forget the actual name at the moment, but one of the Unity Shader tutorials demonstrates this - its the tutorial that displaces vertices of a model face before drawing them). This vertex pass reads from the positions Compute Buffer and modifies the position of the ‘output’ vertex based upon the position read from the input Compute Buffer (used as a StrucutredBuffer in the surface Shader). If it weren’t for this vertex pre-pass,the plat plane would be drawn.

I will show you my code as soon as I’m back at a computer. Good work!

edit added 2nd to last paragraph

1 Like

Okay, back home. i’m going to post my code and provide some explanation ws we go, but don’t have a ton of time at the moment to dive deep into it. Hopefully with the progress you’ve made so far it should be fairly self-explanatory. I’ve tried to remove as much of the ‘non-relevant’ code as possible. Also note that I’m in the midst of a somewhat poor refactoring job, attempting to ‘modularize’ certain aspects of the patch generation. I somehow failed to make a backup of the code prior to this :confused: so it’s not in a compilable state at the moment. That shouldn’t matter - the core code that drives Unity is unchanged, but the organization of it is going to be wrong, so if you see random custom classes referenced in some places but not others, don’t think much of it.

First off, creation of the prototype patch/mesh (dummy mesh). I do this just once, in a class “PatchManager” that is a ‘singleton’ MonoBehaviour.

    private void setupDummyMesh()
    {
        int nVerts = structuralConfiguration.nVerts;
        int nVertsPerEdge = structuralConfiguration.nVertsPerEdge;
        Vector3[] dummyVerts = new Vector3[nVerts];
        Vector2[] uv0 = new Vector2[nVerts];
        int[] triangles = new int[(nVertsPerEdge - 1) * (nVertsPerEdge - 1) * 2 * 3];

        float height = 0;
        for (int r = 0; r < nVertsPerEdge; r++)
        {
            int rowStartID = r * nVertsPerEdge;
            for (int c = 0; c < nVertsPerEdge; c++)
            {
                int vertID = rowStartID + c;
                dummyVerts[vertID] = new Vector3(c, height, r);
                Vector2 uv = new Vector2();
                uv.x = r / (float)(nVertsPerEdge - 1);
                uv.y = c / (float)(nVertsPerEdge - 1);
                uv0[vertID] = uv;
            }
        }

        int triangleIndex = 0;
        for (int r = 0; r < nVertsPerEdge - 1; r++)
        {
            int rowStartID = r * nVertsPerEdge;
            int rowAboveStartID = (r + 1) * nVertsPerEdge;
            for (int c = 0; c < nVertsPerEdge - 1; c++)
            {
                int vertID = rowStartID + c;
                int vertAboveID = rowAboveStartID + c;

                triangles[triangleIndex++] = vertID;
                triangles[triangleIndex++] = vertAboveID;
                triangles[triangleIndex++] = vertAboveID + 1;

                triangles[triangleIndex++] = vertID;
                triangles[triangleIndex++] = vertAboveID + 1;
                triangles[triangleIndex++] = vertID + 1;
            }
        }

        dummyMesh = new Mesh();
        dummyMesh.vertices = dummyVerts;
        dummyMesh.uv = uv0;
        dummyMesh.SetTriangles(triangles, 0);
    }

I also go ahead and set up the material which will be used for later rendering. This is also done only once, in a ‘singleton’ monobehavior. Note that the “ProceduralMeshVertSurf” file must be located in the project’s Resources folder.

    material = new Material(Shader.Find("ProceduralMeshVertSurf"));
    material.SetFloat("_Metallic", 0);
    material.SetFloat("_Glossiness", 0);
    Texture2D texture = (Texture2D)UnityEngine.Resources.Load("GrassRockyAlbedo");
    material.SetTexture("_MainTex", texture);

Now, for each patch I create two ComputeBuffers.

generationConstants, which is where I’ll put the “inputs” to the ComputeShader - i.e. properties required to build the patch like vertex spacing, patch world center, etc.

computeShader.setBuffer(kernel, "terrainGenerationConstants", generationConstants); 

and

patchGeneratedDataBuffer, which the ComputeShader will fill with vertex data (position, normal, etc).

computeShader.setBuffer(kernel, "patchOutput", patchGeneratedDataBuffer); 

Note that the strings above must match those found in the ComputeShader below (I actually do this procedurally but have used plain strings in this example for clarity).

After filling the generationConstants buffer with the appropriate data, I call ComputeShader.setBuffer for the above two buffers, and then dispatch the Compute Shader with

    public void dispatch()
    {
        computeShader.Dispatch(kernel,
        THREADGROUP_SIZE_X,
        THREADGROUP_SIZE_Y,
        THREADGROUP_SIZE_Z);
    }

These constants are defined as:

        public static int nVertsPerEdge { get { return 224; } }     //Should be multiple of 32
        public static int nVerts { get { return nVertsPerEdge * nVertsPerEdge; } }
        public int THREADS_PER_GROUP_X { get { return 32; } }
        public int THREADS_PER_GROUP_Y { get { return 32; } }
        public int THREADGROUP_SIZE_X { get { return nVertsPerEdge / THREADS_PER_GROUP_X; } }
        public int THREADGROUP_SIZE_Y { get { return nVertsPerEdge / THREADS_PER_GROUP_Y; } }
        public int THREADGROUP_SIZE_Z { get { return 1; } }

Here’s the basic version of Compute Shader itself. I’ve left the preprocessor stuff in just to demonstrate a cool technique, but it’s definitely not essential:

#pragma kernel CSMain

#define threadsPerGroup_X 32
#define threadsPerGroup_Y 32
#define nVerticesPerSide 224
#define nVerticesPerSideFloat 224.0

#define TWO_PI 6.283185

#include "noiseSimplex.cginc"
#include "2DNoiseFunctions.cginc"

//#define     RIDGID
#define   HYBRID

#ifdef HYBRID
//Good hybridMultifractal values
#define     NoiseFrequency              0.001
#define     OneMinusFractalIncrement    0.3
#define     Lacunarity                  1.918
#define     nOctaves                    14
#define     MultifractalOffset          0.9

#elif defined RIDGID
//Good ridgedMulti values
#define     NoiseFrequency              0.015
#define     OneMinusFractalIncrement    0.8
#define     Lacunarity                  1.918
#define     nOctaves                    8
#define     MultifractalOffset          0.95
#define     RidgedGain                  1.3
#endif

struct GenerationConstants
{
    float scale;
    float noiseSeaLevel;
    float spacing;
    float4 patchCenter;
};

struct OutputStruct
{
    float4 pos;
};

StructuredBuffer<GenerationConstants> terrainGenerationConstants;
RWStructuredBuffer<OutputStruct> patchOutput;

[numthreads(threadsPerGroup_X,threadsPerGroup_Y,1)]

//We lookup the the index into the flat array by using x + y * x_stride

void CSMain (uint3 id : SV_DispatchThreadID)
{
    GenerationConstants constants = terrainGenerationConstants[0];

    float2 sampleCoord = float2((id.x + constants .patchCenter.x),(id.y + constants .patchCenter.y));

    #ifdef HYBRID
    float noise = hybridMultifractal(NoiseFrequency*sampleCoord, OneMinusFractalIncrement, Lacunarity, nOctaves, MultifractalOffset);
    #elif defined RIDGID
    float noise = ridgedMultifractal(NoiseFrequency*sampleCoord, OneMinusFractalIncrement, Lacunarity, nOctaves, MultifractalOffset, RidgedGain);
    #endif

    float height = constants .scale*max(noise, constants .noiseSeaLevel);

    float4 output = float4(id.x*constants .spacing, height, id.y*constants .spacing, 1);

    int outBuffOffset = id.x + id.y * nVerticesPerSide;
    patchOutput[outBuffOffset].pos = output;
}

So once the ComputeShader completes, outputBuffer will contain the vertex position data. Not depicted here is normals generation, UV coordinates generation, color generation, etc.

Important - I do not try to draw a patch during the same frame in which it was generated. This has caused stalls for me in the past, and I don’t mind waiting a frame. Granted, this was awhile ago when I was using OpenGL (not with Unity), so it might not be a factor here, but I don’t mind waiting a frame to use the patch.

Now that the patch has been generated, on each subsequent frame, and for each and every patch, I do the following:

material.SetBuffer("patchData", patchGeneratedDataBuffer);
Graphics.DrawMesh(dummyMesh, transform.localToWorldMatrix, material, 0, null, 0, null, true, true);

By calling SetBuffer we’re basically hooking-up this patch’s vertex data (generated by the ComputeShader) as an input to the rendering shader. That data will be used to modify the dummyMesh, seen below.

Note that it’s not efficient to call material.SetBuffer like this, because it forces a new Batch to be created (Unity innerworkings). If Unity’s MaterialPropertyBlocks class supported a setBuffer method, that wouldn’t be a problem, but it currently doesn’t. So you’re stuck with 1 patch per batch, which just adds a bit of GPU driver overhead / state thrashing. It’s probably not a huge deal, but definitely not as efficient as it should be.

Inside the "ProceduralMeshVertSurf" shader (loaded and set earlier), we have the following:

Shader "ProceduralMeshVertSurf" {
    Properties {
	_Color ("Color", Color) = (1,1,1,1)
	_colorDeepWater ("Deep Water", Color) = (0.03, 0.16, 0.35, 1.0)
	_MainTex ("Albedo (RGB)", 2D) = "white" {}
	_Glossiness ("Smoothness", Range(0,1)) = 0.5
	_Metallic ("Metallic", Range(0,1)) = 0.0
    }
    SubShader 
    {
	Tags { "RenderType"="Opaque" }
	LOD 200
	
    CGPROGRAM

        #define nVerticesPerSide 224.0
	#define SHOW_GRIDLINES

	#include "UnityCG.cginc"

	// Physically based Standard lighting model, and enable shadows on all light types
	#pragma surface surf Standard fullforwardshadows
	#pragma vertex vert

        struct appdata_full_compute {
            float4 vertex : POSITION;
            float4 tangent : TANGENT;
            float3 normal : NORMAL;
            float4 texcoord : TEXCOORD0;
            float4 texcoord1 : TEXCOORD1;
            float4 texcoord2 : TEXCOORD2;
            float4 texcoord3 : TEXCOORD3;
        #if defined(SHADER_API_XBOX360)
            half4 texcoord4 : TEXCOORD4;
            half4 texcoord5 : TEXCOORD5;
        #endif
            fixed4 color : COLOR;
        #ifdef SHADER_API_D3D11
            uint id: SV_VertexID;
        #endif
        };

	#pragma target 5.0
        sampler2D _MainTex;

	#ifdef SHADER_API_D3D11
	  StructuredBuffer<float4> patchData;
        #endif

        struct Input {
	    float2 uv_MainTex;
	};

        void vert (inout appdata_full_compute v, out Input o) {
            #ifdef SHADER_API_D3D11

            float4 position = patchData[v.id];
            v.vertex = position;

            o.uv_MainTex = v.texcoord.xy;

            #endif
        }

	half _Glossiness;
	half _Metallic;
	fixed4 _Color;

	void surf (Input IN, inout SurfaceOutputStandard o) 
        {
	    #ifdef SHOW_GRIDLINES
	        float2 fract = fmod(IN.uv_MainTex*nVerticesPerSide, float2(1,1));
                fixed4 gridLine = any(step(float2(0.9,0.9), fract));
            #else
                fixed4 gridLine = 0;
            #endif

            #ifdef SHOW_PATCH_BORDER
                fixed4 patchBorder = IN.onBorder;
            #else
                fixed4 patchBorder = 0;
            #endif

            // Terrain color comes from a texture tinted by color
            fixed4 terrainColor = tex2D (_MainTex, IN.uv_MainTex) * _Color;

            fixed4 c = terrainColor + gridLine + patchBorder;

	    o.Albedo = clamp(c.rgb, fixed3(0,0,0), fixed3(1,1,1));

	    // Metallic and smoothness come from slider variables
	    o.Metallic = _Metallic;
	    o.Smoothness = _Glossiness;
	    o.Alpha = c.a;
	}
    ENDCG
    } 
    FallBack Off
}

So the real magic happens under void vert (inout appdata_full_compute v, out Input o) {...}, where the data stored in the patchGeneratedDataBuffer is used to modify the position of the dummyMesh’s vertices.

So there you have it! Or at least the basic process. Happy to answer questions as they come up.

1 Like

Glad to see I wasnt completely on the wrong track in whats going on. Though your strategy seems far more efficient and straightforward, and I wasn’t aware that you can do a Graphics.DrawMesh(…) operation with a dummy mesh.I am right now trying to reimplement this. Suffering from rendering the final plane yet, issue happens while I replace the vertices with

float4 position = patchData[v.id];
v.vertex = position;

in the vertex shader. If I remove these lines the prototype mesh is visible. But using the input buffer for the replacement does not yet work. Expect when I replace the vertice positions within the vertex shader using a noise function (and not the buffer). If I add the vertices from the buffer to the mesh’s vertices by v.vertex += position; , there is some distortion going on.

So I guess the vertex shader is OK but the issue is within reading from the buffer in the vertex shader, or probably more certain within the compute shader. However I am continuing finding the issue. Thanks for this deep insight in your latest post!

Edit: Finally found the issue. It was due to the fact that I didnt size the ComputeBuffer for the output right. I was used to think in Vector3s so I used 12 bytes for the buffer, but didnt realize that I was working on float4 so that I needed to use 16 bytes. But I think its been a good thing to spend 2 days for bugfix searching, so I was forced to to revisit everything and think about what I was doing. Well then, time for the next step, so look how and which shader to best rotate the plane / vertices to their required position/rotation so that I can create a cube with six of those (will review your posts, I am sure you’ve adressed this already).

1 Like

Well getting somewhere. Summary, what do we have right now?

  • A prototype mesh for later replacement in the vertex shader.
  • A constants ComputeBuffer that holds a number of information (nVerticesPerEdge, scale, spacing, patchCenter, planetRadius, noiseSeaLevel) that is being passed to a Compute Shader.
  • A compute shader that creates a plane Z-Up. nVerticesPerEdge * nVerticesPerEdge (so 224x224). Additionally I store all noise results as float in a separate additional buffer (I thought they still could be of use later).
  • A vertex shader that replaces each vertex of the prototype mesh with the
    Note: for each vertex being created in the compute shader I do the following:
    CPU:
    scale = 2.0f / nVertsPerEdge;
    GPU (ComputeShader):
    float4 output = float4(-1+id.xconstants.spacing, -1+id.yconstants.spacing, -1, 1);
    output.xyz = normalize(output.xyz);
    So that each plane is within [-1|1] range and normalizeable. Afterwards the noise is being added.

In order to proceed I have some very short design questions:

  • As right now each patch is created within [-1|1] range and from [-1,-1,-1] to [+1,+1,-1], where are they “shrinked” to take place within the right area (that gets smaller with each quadtree split)? Compute Shader or Vertex Shader?
    Or is the idea simply to decrease the spacing in the constants struct (eventually together with not start rendering from -1+id.x/y)?
  • Are the vertices “brought” to their final position (so basically rotated around (0,0,0) in order to form a quadsphere) within the Compute Shader or within the Vertex Shader?

Of course I could pass over the four edge informations of the quadtree nodes to the generation compute shader and directly place the vertices were they meant to be, but I guess this is not the original idea of instacing the planes and have them all Z-Up first. So it seems there now needs to be some shrinking and rotating being done next.

  • In order to calculate the normals my only idea right now is to have a second pass / second kernel program, “STAGE 2” that runs after all vertices were produced in “STAGE 1”. Which now can check the “surrounding” vertice positions from the buffer. Other less much effective idea is to do additional noise calls. Is there another approach?

  • The surface shader right now uses for testing purposes one texture, _MainTex. In order to have different textures (grass, stone, etc.), is it common style to create one main large atlas texture that holds all kind of terraintypes, or would one better create separate textures per type? Eg

    _GrassTex(“Albedo (RGB)”, 2D) = “white” {}
    _StoneTex(“Albedo (RGB)”, 2D) = “white” {}

2 Likes

Ugh, I can’t tell you how many times this kind of error (“Stride mismatch”) has got me!

Anyhow, congrats! Nice work! You’re quickly out-pacing me (time to quit my day job :slight_smile: )…

So before I address your questions, I need to apologize once again. You have ‘caught up’ to me so quickly that you’re now asking questions about concepts that are not yet incorporated into my “full” Unity project. That project has been been broken (non-compilable) for several weeks as I messed up a large refactoring in an attempt to break things down into multiple classes.

Thus, in order to give you appropriate answers to previous questions, I’ve been digging through an earlier Unity prototype which did not handle an entire planet, only one face of a cube. I’ve also been referring to some even older code from my first planet generator in OpenGL, which rendered an entire planet, but did things a little differently than we’ve been talking about (not large changes, but some subtle ones).

In piecing together the old code with some new concepts, I thought I could ‘smooth over’ the differences. But in reading-back, I’ve noticed several contradictions in my methods. This has probably caused some confusion. So, I must apologize for this in the event that it led you down the wrong path.

Thus In order to answer your most recent questions, I spent a couple of hours today “playing catch-up”, to incorporate the more-complete concepts from my older OpenGL planet and from my design notes into the “full” Unity Project. This project still does not compile, however, due to the botched refactoring I mentioned earlier, and so it remains untested. So the point is: there very well may be bugs in the code.

(I’m considering starting the Unity project from scratch, now that I’ve come up with a better class-design, it might be easier to build up from nothing than continue trying to refactor).

So that having been said, here’s my “actual” ComputeShader (only the relevant parts of the main function), which demonstrates patch generation within the scope of a full planet. I’ve filled in comments to explain the code line-by-line.

void CSMain (uint3 id : SV_DispatchThreadID)
{
      (...)

    // id.x and id.y each run from 0 to nVertsPerSide-1.  id.z == 1 (unused);

    // cubeFaceEastDirection is a vector indicating the cube face's "east" direction
    // cubeFaceNorthDirection is a vector indicating the cube face's "north" direction 
    // Each of these are either (1,0,0), (0,1,0), or (0,0,1), thus each vector is orthogonal to the other.

    // The cube is centered at (0,0,0) and the center of each face is a distance 'planetRadius' from the origin
    // Thus the cube will perfectly inscribe a flat sphere the size of the planet

    // patchCubeCenter is at the center of a quadtree node, sitting on the cube face.

    // 'spacing' is the distance between each vertex in a patch on the face of the cube
    // so for the largest patch (where 1 patch == 1 cube face), spacing === 2*planetRadius / nVertsPerSide
    // spacing is divided in half for each subsequent quadtree subdivision

    // terrainMaxHeight is max height from sea level (i.e. ~ 12 km)

    // first calculate the 'cube space' coordinates of the vertex:

    float eastValue=  id.x - (nVertsPerSide/2.0);
    // eastValue now ranges from -.5*nVertsPerSide to +.5*nVertsPerSide

    eastValue *= constants.spacing
    // eastValue now ranges from -.5*patchWidth to +.5*patchWidth (patchWidth is not actually a defined variable)

    float3 cubeCoordEast = constants.cubeFaceEastDirection * eastValue;

    // now do the same for the "north" direction:
    float northValue=  (id.y - (nVertsPerSide/2.0)) * constants.spacing;
    float3 cubeCoordNorth = constants.cubeFaceNorthDirection * northValue;

    // Now we'll generate the "cube space" vertex coordinate (i.e. restricted to surface of a planet-sized cube)
    // remember, cubeCoordEast and cubeCoordNorth are orthogonal to each other
    float3 cubeCoord = cubeCoordEast + cubeCoordNorth + constants.patchCubeCenter);

    // as an example, if:
    // cubeFaceEastDirection = (1,0,0)
    // cubeFaceNorthDirection = (0,1,0)
    // Then we're dealing with the cube face whose "normal" direction is (0,0,1);
    // Thus patchCenter will, by definition, be somewhere along ( ... , ... , planetRadius)
    // and so:
    // cubeCoord.x will range from [-.5*patchWidth + patchCenter.x, +.5*patchWidth + patchCubeCenter.x]
    // cubeCoord.y will range from [-.5*patchWidth + patchCenter.y, +.5*patchWidth + patchCubeCenter.y]
    // cubeCoord.z will equal planetRadius (for ALL vertices on this patch)

    // To reiterate, at this point 'cubeCoord' is a point restricted to the surface of a planet-sized cube.

    // now we determine the 'planet-space' value for the patchCubeCenter:
    float 3 patchCenter = normalize(constants.patchCubeCenter) * constants.planetRadius;
    // patchCenter now sits on the surface of a planet-sized sphere.

    // next, we normalize the patch
    float3 patchNormalizedCoord = cubeCoord.normalize();  

    // and then calculate its 'real world' size:
    float3 patchCoord =  constants.planetRadius * patchNormalizedCoord;

    // next we 're-center' the patch coordinates to the patchCenter - this is for our 'moving origin' (i.e. camera-centered)
    float3 patchCoordCentered = patchCoord - patchCenter;

    // next we generate the noise value using the patch's 'real-world' coordinate (patchCoord)
    // note: MultifractalOffset (a 3D vector) is used to offset the patch coordinates uniformly, so different planets can be generated.
    // this can be thought of as a different 'input seed' to the noise function.
    // also note: NoiseFrequency is probably << 1.0 (i.e. .0001). This reduces the size of the input coordinates to a size which
    // should not be affected by floating point precision issues
    float noise = hybridMultifractal(NoiseFrequency*patchCoord, OneMinusFractalIncrement, Lacunarity, nOctaves, MultifractalOffset);

    // the hybridMultifractal function (custom function) returns values from 0-1, so they must be scaled to the appropriate terrain height
    float terrainHeight = (noise * 2) - 1;  // terrainHeight now ranges from -1 to + 1;
    terrainHeight *= constants.terrainMaxHeight; // terrainHeight now ranges from -terrainMaxHeight to +terrainMaxHeight. 

    // this final step adds (or subtracts) the real terrain height from the real world-sized (but centered) patch.
    patchCoordCentered += patchNormalizedCoord*terrainHeight;

    int outBuffOffset = id.x + id.y * nVerticesPerSide;
    patchOutput[outBuffOffset].pos = float4(patchCoordCentered.x, patchCoordCentered.y, patchCoordCentered.z, 1);
}

So the end result is a terrain patch that, when translated to “patchCenter” will be located (and oriented) properly in ‘planet-space’ (i.e. planet centered at 0,0,0 world space).

This result is different than what I showed before, and so again, sorry for the confusion. I should have done this ‘translation’ into the new system first, before posting my earlier examples :flushed:

The first question is no longer valid, but the second question is: the vertices are ‘brought’ to their final rotation and scale within the Compute Shader - but the final translation (by patchCenter) must take place in the vertex shader. Again, that is so that the patch can be translated based upon a moving origin (i.e. camera-centered origin).

The vertex shader, then, becomes:

float3 patchCenter;

void vert (inout appdata_full_compute v, out Input o) {
            #ifdef SHADER_API_D3D11

            float4 position = patchData[v.id];
            // translate the patch to its 'planet-space' center:
            position.xyz += patchCenter;

            v.vertex = position;

            o.uv_MainTex = v.texcoord.xy;

            #endif
        }

Of course, note that I am not handling the “moving origin” (nor the location of the planet, for that matter), in the vertex shader. In my old OpenGL planet, I did so - but depending on how you set up Unity and your game objects, you can perform these translations entirely on the CPU. The surface shader applies all of the appropriate gameObject transformation “Behind the scenes” (somewhere in between the vert function and the surf function) according to the GameObject transforms. CAVEAT: to avoid precision issues, you’ll need to store camera, planet, and patch locations as doubles, then do the appropriate subtractions as doubles and finally down-sample to floats before passing the ‘final’ translation on the the GPU. And I believe Unity only stores the transforms as floats…so you may have to manage your own double-precision transformation on the CPU. I haven’t gotten this far with Unity, so I can’t give details.

But the reason that the patch is stored in the compute buffer ‘re-centered’ to the ‘patchCenter’ is to reduce the sizes of the floating point numbers. If you store the patch vertices as ‘real world’ positions with an origin of (0,0,0), (i.e. the vertices are planetRadius +/- maxTerrainHeight from the origin), then you’ll experience floating-point precision errors. Instead, by using the ‘patchCenter’ as the origin, then the vertices are only (at most) sqrt(2*(patchHalfWidth*patchHalfWidth) + terrainMaxHeight*terrainMaxHeight) from the origin.

In my OpenGL code, I performed “STAGE 2” as you describe, using transform feedback (a less-flexible type of Compute Shader) to numerically calculate the (approximate) derivatives (slope) of the terrain patch. It turns out you may not need to do this, however, because you can get the derivative of a simplex noise function analytically. See this: http://staffwww.itn.liu.se/~stegu/simplexnoise/DSOnoises.html …But, the math might get a bit complex, since you’d need to factor-in the curvature and orientation of the terrain patch. I have not yet done this, and don’t know if it would be more efficient than running a second pass (but my guess is that it would be more efficient).

I haven’t gotten that far with the Unity project - I’m also using just one Texture. (And BTW, my texture coordinates aren’t correct, either, because they don’t wrap properly at face edges - I’m considering using tri-planar mapping, but that approach is a bit expensive and can produce ‘blurriness’ when blending between textures. That’s on my to-do list. :confused: ) But to your question, atlases are more efficient in general, but I’m not sure how Unity’s back-end handles them. Plus if your textures are large you might exceed the GPU’s max texture size by packing them into an atlas.

With regard to selecting the appropriate texture, I believe the ‘material type’ should be stored on a per-vertex basis - at least, this is what I did in my OpenGL project. I determined the appropriate texture type in STAGE 2, since I had slope information in that stage, and used that information to select a different texture (i.e. rocky texture if slope is greater than a certain value).

Long post. Hope it’s helpful. Final apologies for earlier confusion. From here forward, I’ll refer only to my ‘current’ approach to avoid mixing concepts as I did earlier. But by the looks of it, soon enough you’ll be helping me!

Navy

Hi NavyFish,

thank you for the update. I implemented the new strategy ~2-3 weeks ago, unfortunately I did a “code-all-at-once-and-test-afterwards” approach and it didnt work :yum: . Well could’ve been anything, so today I started again building up on my working implementation, changing and checking everything step by step.

And yes, the extended constants buffer and compute shader, following your strategy, now works :grinning: . Very nice. Looking at it, it should be straightforward to append this approach to the six quadtrees for creating a sphere, as in the end it seems it is all about changing the parameters in the constantsbuffer per quadtreenode, the rest is done in the computer shader.

There is one issue left right now, as the generated mesh only reflects light when the light source is behind the plane, as seen on the screenshot. Not sure why this happens, but I havent yet implemented normals calculation, so we’ll see. But as soon as I fixed this, the plan is to append the plane creation in this shader to the six quadtree nodes next, to have a sphere I can work on.

2 Likes

Since you made it to shaders, here is a Logarithmic Depth Buffer implementation for Unity, made from Outerra blog post, might come in useful. There is the issue that Unity Surf Shader doesn’t accept a depth value as input, so the implementation is a Frag one instead of Surf.

Anyway…
http://pastebin.com/FJfEnWy7

Also, for testing normals, just normalize your vertex position if the planet is centered at zero, that will give you a perfect sphere’s normals.

That would give normals for a sphere, yes, but wouldn’t account for any terrain features.

Regarding lighting behind the plane, it sounds like the triangle winding order might be reversed. Hard to say for sure, not super familiar on Unity’s requirements.

Great work! I’m writing from the road so forgive the short response. Looking forward to tracking your continued progress!

Almost there. The lightningt now works after I implemented the normal calculation in the second pass / kernel in the compute shader. For the moment I did the simple approach (just use normalized coordinates for spherical lightning), as soon as the full sphere is generated I’ll be continuing in implementing the more accurate approach calculating the normals on their neightbours to take the noise into account.

I currently work on creating multiple planes (6, for a sphere). Right now I try to figure out if there is some dependency on when the buffers and materials are being created, the DrawMesh call done and if I need a Unity3D gameobject (as right now no two meshes are rendered at once) per DrawMesh call.
So right now my (not working) approach is I moved the buffers into the QuadtreeTerrain class (a quadtree node), as well as the material (not sure if individual materials are necessary).

class QuadtreeTerrain {
    // Quadtree classes
    public QuadtreeTerrain parentNode; // The parent quadtree node
    public QuadtreeTerrain childNode1; // A children quadtree node
    public QuadtreeTerrain childNode2; // A children quadtree node
    public QuadtreeTerrain childNode3; // A children quadtree node
    public QuadtreeTerrain childNode4; // A children quadtree node
    // Buffer
    public ComputeBuffer generationConstantsBuffer;
    public ComputeBuffer patchGeneratedDataBuffer;
    // Material
    public Material material;
        ....
}

In the SpaceObjectProceduralPlanet script, applied to a single game object, I hold six instances of quadtrees [=QuadtreeTerrain] then.

public class SpaceObjectProceduralPlanet : MonoBehaviour {
    ....
    // QuadtreeTerrain
    private QuadtreeTerrain quadtreeTerrain1;
    private QuadtreeTerrain quadtreeTerrain2;
    private QuadtreeTerrain quadtreeTerrain3;
    private QuadtreeTerrain quadtreeTerrain4;
    private QuadtreeTerrain quadtreeTerrain5;
    private QuadtreeTerrain quadtreeTerrain6;

    // We initialize the buffers and the material used to draw.
    void Start()
    {
        ...
        // QuadtreeTerrain
        this.quadtreeTerrain1 = new QuadtreeTerrain(0, edgeVector1, edgeVector2, edgeVector3, edgeVector4, quadtreeTerrainParameter1);
        this.quadtreeTerrain2 = new QuadtreeTerrain(0, edgeVector2, edgeVector5, edgeVector4, edgeVector7, quadtreeTerrainParameter2);
        this.quadtreeTerrain3 = new QuadtreeTerrain(0, edgeVector5, edgeVector6, edgeVector7, edgeVector8, quadtreeTerrainParameter3);
        this.quadtreeTerrain4 = new QuadtreeTerrain(0, edgeVector6, edgeVector1, edgeVector8, edgeVector3, quadtreeTerrainParameter4);
        this.quadtreeTerrain5 = new QuadtreeTerrain(0, edgeVector6, edgeVector5, edgeVector1, edgeVector2, quadtreeTerrainParameter5);
        this.quadtreeTerrain6 = new QuadtreeTerrain(0, edgeVector3, edgeVector4, edgeVector8, edgeVector7, quadtreeTerrainParameter6);
        CreateBuffers(this.quadtreeTerrain1);
        CreateBuffers(this.quadtreeTerrain2);
        CreateBuffers(this.quadtreeTerrain3);
        CreateBuffers(this.quadtreeTerrain4);
        CreateBuffers(this.quadtreeTerrain5);
        CreateBuffers(this.quadtreeTerrain6);
        CreateMaterial(this.quadtreeTerrain1);
        CreateMaterial(this.quadtreeTerrain2);
        CreateMaterial(this.quadtreeTerrain3);
        CreateMaterial(this.quadtreeTerrain4);
        CreateMaterial(this.quadtreeTerrain5);
        CreateMaterial(this.quadtreeTerrain6);
        Dispatch(this.quadtreeTerrain1);
        Dispatch(this.quadtreeTerrain2);
        Dispatch(this.quadtreeTerrain3);
        Dispatch(this.quadtreeTerrain4);
        Dispatch(this.quadtreeTerrain5);
        Dispatch(this.quadtreeTerrain6);
}

    // We compute the buffers.
    void CreateBuffers(QuadtreeTerrain quadtreeTerrain)
    {
        .... preparing generation constants
        quadtreeTerrain.generationConstantsBuffer.SetData(generationConstants);
        // Buffer Output
        quadtreeTerrain.patchGeneratedDataBuffer = new ComputeBuffer(nVerts, 16 + 12 + 4 + 12); 
    }

    //We create the material
    void CreateMaterial(QuadtreeTerrain quadtreeTerrain)
    {
        Material material = new Material(shader);
        material.SetTexture("_MainTex", this.texture);
        material.SetFloat("_Metallic", 0);
        material.SetFloat("_Glossiness", 0);
        quadtreeTerrain.material = material;
    }

    //We dispatch threads of our CSMain1 and CSMain2 kernels.
    void Dispatch(QuadtreeTerrain quadtreeTerrain)
    {
        // Set Buffers
        computeShader.SetBuffer(_kernel, "generationConstantsBuffer", quadtreeTerrain.generationConstantsBuffer);
        computeShader.SetBuffer(_kernel, "patchGeneratedDataBuffer", quadtreeTerrain.patchGeneratedDataBuffer);
        // Dispatch first kernel
        _kernel = computeShader.FindKernel("CSMain1");
           computeShader.Dispatch(_kernel, THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
        // Dispatch second kernel
        _kernel = computeShader.FindKernel("CSMain2");
        computeShader.Dispatch(_kernel, THREADGROUP_SIZE_X, THREADGROUP_SIZE_Y, THREADGROUP_SIZE_Z);
    }

    // We set the material before drawing and call DrawMesh on OnRenderObject
    void OnRenderObject()
    {
        this.quadtreeTerrain1.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain1.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain1.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain2.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain2.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain2.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain3.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain3.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain3.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain4.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain4.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain4.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain5.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain5.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain5.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);

        this.quadtreeTerrain6.material.SetBuffer("patchGeneratedDataBuffer", this.quadtreeTerrain6.patchGeneratedDataBuffer);
        Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, this.quadtreeTerrain6.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);
    }

    //When this GameObject is disabled we must release the buffers.
    private void OnDisable()
    {
        ReleaseBuffer();
    }

    //Release buffers and destroy the material when play has been stopped.
    void ReleaseBuffer()
    {
        // Destroy everything recursive in the quadtrees.
        this.quadtreeTerrain1.generationConstantsBuffer.Release();
        this.quadtreeTerrain1.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain2.generationConstantsBuffer.Release();
        this.quadtreeTerrain2.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain3.generationConstantsBuffer.Release();
        this.quadtreeTerrain3.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain4.generationConstantsBuffer.Release();
        this.quadtreeTerrain4.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain5.generationConstantsBuffer.Release();
        this.quadtreeTerrain5.patchGeneratedDataBuffer.Release();
        this.quadtreeTerrain6.generationConstantsBuffer.Release();
        this.quadtreeTerrain6.patchGeneratedDataBuffer.Release();
        DestroyImmediate(this.quadtreeTerrain1.material);
        DestroyImmediate(this.quadtreeTerrain2.material);
        DestroyImmediate(this.quadtreeTerrain3.material);
        DestroyImmediate(this.quadtreeTerrain4.material);
        DestroyImmediate(this.quadtreeTerrain5.material);
        DestroyImmediate(this.quadtreeTerrain6.material);
    }

    void Update() {
        // Do nothing
    }

}

Of course this is very bruteforce, but well this should work before I proceed as I need to figure out how to handle the buffers and draw calls and where to put them. I am a bit afraid I could need one gameobject per DrawMesh call, because I was hoping I could avoid multiple gameobjects (it may be that it could make sense anyway when I want to continue with colliders sight). Will investigate on that and edit this post then as soon as this is fixed. Then the fun stuff (refine the noise for a planet like visual and texture it) should come next, as well as scaling things up to 1:1 scale (right now I am on 1:1000 for debug purposes in the editor).

1 Like