Procedural Terrain Rendering How-To

Sorry I dont exactly know what you are doing and how, but IMHO

a) the X and Y coordinates of PatchCubeCenter should be mostly positive for ID0 - Bottom Right and ID1 - Top Right (they arent in your examples)
b) I newer change CubeFaceEastDirection and CubeFaceNorthDirection once I start splitting (why should you, they remain on the same side of the cube).

Hm, no. PCC need to be shifted as LOD level increases.
Look at this… This is representation of ZPositive quad 2lvl LOD, Front quad of a planet.


So, i don’t know how to explain in an other words, cuz my native language is russian :smile:

Edit: Only @NavyFish can explain dat operations))))
On my CPU generator i calculate corner positions for each surface(quad) and use it when i need to subdivide/merge…

CubeFaceEastDircection and CubeFaceNorthDirection remain the same for each split, because they are on the same side of the cube.

No. The PCC remains always the same for each quadtree node instance. What you have to do is to create new center coordinates for the 4 child node instances (based on their new outer vectors calculated out of those of the parent). But you never literally shift a center (or anything else) of a node instance in your tree. You “calculate” for each child new four outervector coordinates based on the parents one (easy, four of them are already there in your parent, only the five “half” new outervectors need to be calulcated). No need to care about the centervector specifically (you have it already when you have the outer ones, as its just the center of the four, I was long time curious if I even should save it in a class variable instead of returning it by a small function in the node).
Thats the beauty of the tree, by splitting a leaf node and considering the outervectors of the parent during creation of the child node (with new outer vectors) the tree manages itself. You dont reuse a quadtree node for being something else nor you shift, you create (or delete in case of a merge) new child instances each time, each with its own variables (and as I did own buffers), representing in case of a procedural planet a specific plane at a specific coordinate. To be more drastic, in the end each parent plane and its (in case of LOD change) child planes (all four together covering the same area as the parent plane!) have nothing todo with each other except that there is a child<->parent relation in your quadtree and that the children consider the outervectors of its parent (and as you know that you might reuse some information of the parents, e.g. noise buffers etc.).

(quickly drawn picture, havent checked if all vector coordinates are correct, no time)

Maybe I am all wrong, thats my understanding how to use the quadtree and how I go with my approach. Maybe NavyFish or cybercritics have another quadtree approach and I am the only one doing it that way :relaxed:

Wow, I never expected to get such a detailed, and helpful answer. If you ever visit Iceland then I’ll make sure to invite you for a hot cup of coffee! :wink:

It’s actually quite amazing what you can do with WebGL these days, especially if you abuse the shaders! I really recommend checking out http://shadertoy.com to get an idea what people are cooking with shaders alone.

Regarding your detailed answer, I’m still digesting all the information, especially the part that I can just forget the Jacobian matrix. I also like the part that you think about optimization all the way, something that I can spend hours playing with. E.g. using squared distance check to save a few sqrts!

Since I’m talking about optimization and square roots, I recommend checking out this thread that I started last year:
http://forum.outerra.com/index.php?topic=3324.0

cameni shared his cube to sphere mapping algorithm that only uses one sqrt. The classic solution from Philip Nowell uses three sqrts. It’s amazing how much cameni has contributed to the world of procedural planets. Hopefully we’ll be seeing more golden information from him, and as well from others (like you!).

Now back to your answer. I’ve computed the normals as you suggested, and the normal map “seems” to look right. Now, for each patch, the center of it is defined in world space, and the rest of the vertices in a patch are relative to the center. I subtract the camera position from the center position of a patch before its sent to the GPU (to avoid precision jumping/shaking when I’m close to the surface at LOD level 20+).

I’m still a bit confused on how the TBN vectors come into play. Vobj is a vertex on the patch, right? Which should be in patch space, or at least relative to the center of the patch. I shouldn’t need to do any multiplications if that’s the case, or at least that’s how I understand it. If not, do I just need to multiply the vertex position with the BTN matrix to get it into patch space (if the vertex is in world space)? It would be awesome if you could go slightly into this topic as if you were explaining it to someone who is very rusty in matrices! :smiley:

Thank you again NavyFish, I really appreciate the time you’ve taken to answer me, and others here. Hopefully I can contribute something back soon!

*edit1: How are you accounting for the patch curvature in the normal calculation? When I think of it, I could simply store the sphere normal on each vertex and interpolate it in the fragment shader? I feel like I’m missing something :slight_smile:

*edit2: Doing my homework, sometimes it’s just so addicting to ask more questions! - http://www.opengl-tutorial.org/intermediate-tutorials/tutorial-13-normal-mapping/ - explains everything pretty well!

So. After spending two days to figure out the @NavyFish’s formula for shifting PCC after splitting - result is fail.
For example repositioning method (Only for LOD level 2) is 445 lines wide, and code count = exp(LODLevel) :smile:

Btw, after second attempt, when i do some corrections on shader - fail too…

Now i gonna cleanup my GPU generator, and make the same calculations on the GPU, if splitting method will be not researched :smiley:

Btw, i already have configurable noise engine based on my CPU algorithms, which producing pretty results.

Edit: I have small progres with quad corners…

Edit: Looks like i missed smth…

Edit: So. I get it. Up to 6 lvl’s of lod. Yaahuu!
Edit: Seems my system works good with a custom count of the subquads per parent quad…

Has anyone of you implemented a good Frustum Culling strategy in Unity3D yet?

As I right now only use horizon culling, frustum culling is still to be done but rather important! My previous first version was using WorldToScreenPoint (and checking if Z is positive or negative), but I am not satisfied with this.

Right now I am thinking of adding to my quadtree-parsing a similiar method I used to check if the camera is within planetcenter-to-plane volume of the plane (to check if I need to skip horizon culling).

(the above picture shows how to determine the shortest length to a plane at a spherical sphere, but its first step is the determination if the camera is inside or outside the volume)

The code is simple, its a test four times runned for each side of the plane, using one dot and one cross function to check if the camera position is on the right side (inside or outside) per edge. If for all four tests the position of the camera is on the inner side of the plane’s edge, the point is inside the volume.

For frustum culling I could use this test the other way round and get the six planes of the camera’s frustum using the GeometryUtility.CalculateFrustumPlanes function once per frame before parsing all six quadtrees, and then check for each plane edge if one of the four edges is within the frustum volume (if yes, no culling).

However I wonder if one did implement a different frustum approaches while parsing the quadtree and trying culling.

This actually sounds like a brilliant way! Just to clarify one thing, when you speak of patch space, has the mesh been mapped to sphere in the patch space? And if so, how does the biTangent = cross(normal, float3(1,0,0)); work if the patch mesh is slightly curved? Like for example if you look at the lowest LOD patch, the curve is quite steep, wouldn’t the biTangent be incorrect on the patch edges?

Further explanation would be awesome, because this method would make so many calculations easy as pie!

Thanks.

Great question… curvature won’t affect things, but you will need to normalize the result of the cross products if there’s curvature.

This would fail if at any point your normal equaled or was very very close to (1,0,0). If that were the case, you could simply cross the normal with float3(-1,0,0), then negate the result. Checking for that error case should be fairly cheap: assuming your normals are always normalized (they should be!), then you just check if 1 - abs(normal.x) < .01

Of course, you want to reduce branching as much as possible in shaders, so I’d actually write that like:

float epsilon = .01; //define this as a constant or use a pre-processer #def for it
float refVectorSign = sign(1 - abs(normal.x) - epsilon);
vec3 refVector = refVectorSign * float3(1.0, 0, 0);

vec3 biTangent = refVectorSign*cross(normal, refVector);
vec3 tangent = cross(-normal, biTangent);

My patches are defined in patchSpace, and do already incorporate curvature, which is determined during my first stage (vertex position generation, as opposed to the 2nd stage where normals are calculated)… Normal calculation works perfectly fine with that curvature in-place. As far as the algorithm is concerned, height differences caused by the noise functions are exactly the same as height differences caused by curvature. Hope that helps.

@zameran

I want to help you but am not totally sure what’s wrong. Perhaps if you shared your entire quadtree code it would help?

But from what I can tell, it looks like your’re having difficulty with the quadtree structure itself. For that I would recommend reading a tutorial or two about quadtrees. This one is pretty good: http://www.gamedev.net/page/resources/_/technical/graphics-programming-and-theory/quadtrees-r1303

Also - please realize, we’re doing the quadtree on the CPU. The vertices of the terrain mesh are generated on the GPU, but the CPU is responsible for setting up (i.e. splitting and merging) the quadtree.

Sorry I can’t help more than that… check out a few tutorials (I’m sure you can find a quadtree tutorial in Russian too) on quadtrees, that might answer your question. And if not, come back and share more of your code.

@JoergZdarsky

Here’s an answer that will make you smile: Unity does frustum culling for you!

As long as your mesh’s bounds are correctly set, Unity will cull whatever is outside the camera’s frustum.

Remember when your planet was disappearing prematurely? That was unity performing its frustum culling… except with wrong bounds.

EDIT Check out this event: http://docs.unity3d.com/ScriptReference/Camera.OnWillRenderObject.html

Use that to count how many of your patches are being rendered (and thus to determine how many aren’t)

Ok. I can share my code, but can i do it on PM? Thank you.

@JoergZdarsky

I have inplemented simple Frustum culling in my CPU project. So it does not perfect, but it can turn you in a good direction.

BorderFrustumCheck Method:

public bool BorderFrustumCheck(Camera camera, Vector3 border)
{
    float offset = 64.0f; //128
    float disturbtionOffset = 8.0f;
    bool useOffset = true;
    bool useDisturbtion = false;
    Plane[] planes = GeometryUtility.CalculateFrustumPlanes(camera);
    for (int i = 0; i < planes.Length; i++)
    {
        if (useDisturbtion)
        {
            for (int j = (int)(offset - disturbtionOffset); j < offset + disturbtionOffset; j++)
            {
                if (planes[i].GetDistanceToPoint(transform.TransformPoint(border)) <= (useOffset ? 0 - offset + j : 0 + j))
                {
                    return false;
                }
            }
        }
        else
        {
            if (planes[i].GetDistanceToPoint(transform.TransformPoint(border)) <= (useOffset ? 0 - offset : 0))
            {
                return false;
            }
        }
    }
    return true;
}

PlaneFrustumCheck method:

public bool PlaneFrustumCheck(Camera camera, Surface surface)
{
    if (BorderFrustumCheck(camera, surface.topLeftCorner) ||
        BorderFrustumCheck(camera, surface.topRightCorner) ||
        BorderFrustumCheck(camera, surface.bottomLeftCorner) ||
        BorderFrustumCheck(camera, surface.bottomRightCorner))
    {
        return true;
    }
    else
        return false;
}

Screenshot:
Big size!

So, I say it again - it is very simple inplementation, and it works good for me. Realy. I have a lot of culling methods inplemented, and i switching these.
E.g On far distances to planet i’am using Angle Culling. (Simple check angle of surface middle point normal to camera), On middle distances i’am using Plane Volume Culling.
From atmosphere maximum height to ground i’am using Frustum Culling.
I can record a video of “How it works, and how it looks like in an action” - just say.

.

OK just tested if this approach with our setup where there is only a single mesh prototype instance. And in fact it works when setting the boundaries before calling Graphics.DrawMesh during the OnRenderObject().

    /// <summary>
    /// Called to render the planet's plane each frame. It sets the material and calls the shader to draw all elements in the QuadtreeTerrainRenderQueue
    /// </summary>
    void OnRenderObject()
    {
        // Render
        for (int i = 0; i < this.QuadtreeTerrainRenderQueue.Count; i++)
        {
            QuadtreeTerrain quadtreeTerrain = this.QuadtreeTerrainRenderQueue[i];
            if (quadtreeTerrain.quadtreeTerrainState == QuadtreeTerrain.QuadtreeTerrainState.READY)
            {
                // Set Bounds
                this.prototypeMesh.bounds = new Bounds(quadtreeTerrain.WScenterVector, quadtreeTerrain.WSbounds);
                // Set Material
                quadtreeTerrain.material.SetBuffer("patchGeneratedFinalDataBuffer", quadtreeTerrain.patchGeneratedFinalDataBuffer);
                quadtreeTerrain.material.SetTexture("_MainTex", quadtreeTerrain.patchGeneratedSurfaceMapTexture);
                quadtreeTerrain.material.SetTexture("_NormalMap", quadtreeTerrain.patchGeneratedNormalMapTexture);
                // Set Pass
                //quadtreeTerrain.material.SetPass(0);
                // Draw Mesh
                Graphics.DrawMesh(this.prototypeMesh, transform.localToWorldMatrix, quadtreeTerrain.material, LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName), null, 0, null, true, true);
            }
        }
    }

Interesting. It helped me to push the triangles significantly down to ~200k to 2.0M triangles depending on the distance.
I tested it if would be more efficent to create a prototype mesh per quadtree node so that I do not need to set the bounds each frame but once - but it wasnt, the above version was significantly better especially at close-range.
Unity starts to cull too early at higher LOD, but thats probably an issue somewhere I still need to Debug why. But anyway visually and performance-wise a massive improvement, combined with now that I’ve rebuilt the shaders and switched fully to a highres normalmap (128x128) and surface texture (128x128) besides a lowres vertex mesh (32x32).

The issue now is the CPU performance when getting deeper into the quadtrees, at ~9-10 it drops to significantly low to 50-100ms per frame and worse. So thats next part of optimization which requires profiling where I am loosing time.

  • Scan of all six quadtrees per frame
  • The split or merge process incl. the compute shader dispatches

Thanks @zameran I am going to take a look at your culling methods. Yes I was thinking about switching culling methods too, as I suspect the depth of the quadtrees right now to be the issue.

Of course videos are always good and if its just for eyecandy for everyone! :smile:

Ok, @JoergZdarsky i will record some videos, and post it here.

About

The issue now is the CPU performance when getting deeper into the quadtrees, at ~9-10 it drops to significantly low to 50-100ms per frame and worse. So thats next part of optimization which requires profiling where I am loosing time.

I have this on CPU Generator, cuz i’am using a lot of methods in my surfaces, more surfaces(higher LOD level) - more m/s.

So. Very easy to fix - use “queue array” to UpdateLOD/Generate/Dispatch/GenerateCollider/etc etc.
And do it multithreaded.
Guys can you push me in true direction with generation and calculation of “fake mesh” (wich vertCountPerSide bigger than normal mesh) for true calculation of normals/CPUErosion/other…
My first attempt was failed…

Thanks @zameran I am going to take a look at your culling methods. Yes I was thinking about switching culling methods too, as I suspect the depth of the quadtrees right now to be the issue.

Should i share all my culling merhods, that is already in use now on CPU generator? With explanations…

Have you read posts #113 and #116? Procedural Terrain Rendering How-To
Thats my exact code for this process, where nVerts == 226.

My second stage only dispatches 224x224 threads, and I calculate the array indexes so that the kernel does not encounter vertices on the edge (I.e. it runs from 1-225 in both dimensions). The normal is then calculated as discussed recently.

I’m out skiing today and tomorrow so won’t have access to my code until returning home Sunday evening. Hope that helps though.

Nice frustum culling techniques! I like the three different approaches.

Ya. I see that. To say in truth - my code is the same :slight_smile: So.
When i trying to “clamp” dispatch id - output is mess…

Could you show me the example?

I’m out skiing today and tomorrow so won’t have access to my code until returning home Sunday evening. Hope that helps though.

Have a nice day!

Nice frustum culling techniques! I like the three different approaches.

I like too!

Edit: My friend ask me for a stream… I make it.
RU_Recording (1440x900)
So idk how to make different sized calculation’s - can’t understand several things.

I realy wanna recode all, with obtained experience.

Thank you again for the great explanation. My normal map in tangent space seems to be working beautifully, with one exception. See screenshots below:

Far away:

High above the surface, ignore the cracks, I will fix them later:

With each LOD level the normal map seems to fade out, as if the normal strength gets weaker with each LOD level.

I get better result if I increase the normal strength with each LOD level. But is this a normal behaviour, or a bug with my normal calculations?

Instead of hard-coding the strength to 1.0 / 18.0 I divide by depth (LOD level).

vec4 N = vec4(normalize(vec3(dX, dY, 1.0 / 18.0)), 1.0);

Is changed to:

vec4 N = vec4(normalize(vec3(dX, dY, 1.0 / depth)), 1.0);

What do you guys think?

Definitely a bug. Your vertexSpacing value is incorrect. It needs to be the actual vertex spacing at each LOD, otherwise the normal will not be built correctly.

float3 normal = normalize(float3(dx, dy, 2*vertexSpacing))

Vertex spacing should go as a function of LOD as such:

vertexSpacing = spacingAtLODZero / (LOD ^ 2)

Where LOD starts at 0 (largest patches) and increases from there. In other words, vertex spacing should decrease by half each LOD.

This explains your image. By not decreasing vertexSpacing at each LOD, the vertical component of your normal vector becomes relatively larger and larger than what it should be, compared to the scale of the horizontal components. If you were to depict your normals visually, I bet they’d all be pointing straight up (or close to it). Do you have the appearance of a strong ‘specular’ effect (almost like reflective) when you look straight down at the surface?

@zameran Here’s how I calculate my normals:

[The following two shaders are dispatched with thread group sizes of (1, 1, 1) (although eventually this will change)]

First stage, generation of vertex positions:

#define nVertsPerSideWithBorder 130

[numthreads(nVertsPerSideWithBorder , nVertsPerSideWithBorder , 1)]
void CSMain (uint3 id : SV_DispatchThreadID)
{
    int outBuffOffset = id.x + id.y * nVertsPerSideWithBorder;

    //generate height value, i.e.
    float height = amplitude*sin(frequency*TWO_PI*(id.x - offsetX)/(float)nVertsPerSideWithBorder );
    ...

    positionsOutput[outBuffOffset].pos = float4(id.x*spacing, id.y*spacing, height, 1);
}

Second stage, generation of normals:

#define nVertsPerSideWithBorder 130
#define nVertsPerSide 128

[numthreads(nVertsPerSide , nVertsPerSide , 1)]
void CSMain (uint3 id : SV_DispatchThreadID)
{
    int inBuffOffset = (id.x + 1) + (id.y+1) * nVertsPerSideWithBorder;  //'input' index of positions buffer
    int outBuffOffset = id.x + id.y * nVertsPerSideWithBorder;               //'output' index of vertexBuffer

    float n = positionsInput[inBuffOffset + nVertsPerSideWithBorder].z; 
    float s = positionsInput[inBuffOffset - nVertsPerSideWithBorder].z;
    float e = positionsInput[inBuffOffset + 1].z;
    float w = positionsInput[inBuffOffset - 1].z;
    
    float ew = e-w;
    float ns = n-s;
    float3 result = normalize(ew, ns, 2*spacing);

    vertexOutput[outBuffOffset].normal = result;  //store normals in 128x128 output buffer
    vertexOutput[outBuffOffset].position = positionsInput[inBuffOffset].pos;   //copy vertex positions from 130x130 input buffer to 128x128 output buffer
}

I hope that helps!

You’re right, my vertex spacing was incorrect. It helped when I corrected it, but it’s still weak/faded out when I get closer to the surface.

I manually calculated the spacing between vertices on LOD level 0, which is 0.125.

float vertexSpacing = 0.125 / ( (depth + 1.0) * (depth + 1.0) );
vec3 normal = normalize(vec3(dx, dy, 2.0 * vertexSpacing));

The variable depth is 0 at LOD level 0, so I added 1.0 to it to avoid division by zero. I’ve tried playing around with the variables, but it’s never “perfect”.

Here’s a video I recorded:

Maybe I shouldn’t worry too much about it at this point.

Thanks!

try:

float vertexSpacing = 0.125 / pow(2, depth);

when depth == 0, pow(2, depth) ==1. No division by zero.

By the way, it looks nice! What language are you using? Is that your own engine?

Awesome, that did the trick! It’s 100% seamless now! :slight_smile: I competely forgot about pow().

Thanks! It’s written in pure Javascript and GLSL (OpenES 2.0), and yes that’s my own engine. I’ve thought about porting it to some javascript WebGL engine like Three.js, but I really like doing everything from scratch (a bad habbit!).

Thanks again for the help! How’s your project coming along?

Edit1: Video to enjoy!

Edit2: Artefacts start to show up around LOD level 11. There are probably a few things that need to be fixed. E.g. moving the vertex spacing to the CPU for more precision?