Procedural Terrain Rendering How-To

Well not really procedural coding news from my side, but some really great news for us Unity users.
Too bad I have hardly time for coding left currently, but this experimental implementation will be for sure the next thing I am going to try out. Could be the solution we’ve been waiting for for performance and good culling mechanism.


Watch at 1:33:19

Already added as experimental API in 2018.1b:

AsyncGPUReadback
class in UnityEngine.Experimental.Rendering/

Implemented in: UnityEngine.CoreModule
Description: Allows the asynchronous read back of GPU resources.

This class is used to copy resource data from the GPU to the CPU without any stall (GPU or CPU), but adds a few frames of latency. See Also: AsyncGPUReadbackRequest.

https://docs.unity3d.com/2018.1/Documentation/ScriptReference/Experimental.Rendering.AsyncGPUReadback.html
https://docs.unity3d.com/2018.1/Documentation/ScriptReference/Experimental.Rendering.AsyncGPUReadback.Request.html

2 Likes

flagged as spam

5 Likes

NECRO!!!

Looking forward to early access release.

Anyone done anything cool lately? Sadly I haven’t written a line of code in a couple years now. But, it comes and it goes.

Hope everyone is doing well, and I look forward to seeing some of you in game in the near future!

Navy

7 Likes

Hello, everybody,

is anyone else working on a planet?

I’ve just upgraded to D3D11, here’s a short video.

3 Likes

Working on Texturing now (for debug just 3 Color,…)


Have some issue at the Ground, need to figure out how i get more Details on ground.

3 Likes

Worked on a new revision of my instantiation and placement of stars in Unity3D in real scale. Based on some techniques I described in this thread, but now the positioning of stars completely bases on my custom coordinate system and recalculates the position of objects each frame now new from scratch. Came now up with something I am pretty satisfied with, implemented the freelook similar to „Space Engine“, and now move forward to implement cubesphere based bodies again and remove the debugging meshes. The video does not show any eyecandy :slight_smile:, but the performance involved some work :-).
PS: glad everyone is fine!

EDIT: Oh my I hate it when you forget stuff you worked on previously. I started implementing the quadtree structure again and now consider the advice to create all planes Y-up model space (up being [0,1,0] first for AABB testing and then rotate the transform to the desired plane location (which would be the center stored in a node)., e.g. [0,0,-1] which would be the back plane for the cubesphere.
Now I stumble across the whole azimuth and elevation rotation on how to rotate each Y-up plane to its location. My old code where I rotated planes to Y-up (where I now have to do the oposite) wasnt well enough document, as my current testing gives weird results, ARGH.

5 Likes

The Only Boundary is Your Imagination…

2 Likes

Hi everyone,

long time no planetary terrain update in one of the best I-Novae threads :-). I am currenty revisiting planetary bodies using cubespheres and quadtrees, and thought I might revive this thread for some new “best strategies” or “common approaches”. While currently re-implementing (in current Unity3D 2019 version) I want consider hints in some previous comments in this thread, but already noticed they raise some questions.
Idea is to aks for critics, hints, corrections, different approaches.

Lets lets try to structure the topc into quadtrees, terrain-plane pipeline and maybe some CPU vs GPU discussion.

QuadTree Structure

  • Each quadtree deals with one plane of the non-normalized cube.
  • Each cubesphere has six independend quadtrees. Each quadtree manages one side of the cube.
  • A rendered plane (=its vertice data etc) is available (only) in the leafnodes of the quadtree. Once we split a node, we destroy its current plane data (including vertices) and calculate the new plane data in realtime in the 4 childsnodes (leafs).
  • We traverse each of the 4 quadtree recursively completely and check if we have to split or merge quadtree nodes based our LOD strategies.
  • The very basic components of a quadtree node are:

public interface IQuadTreeNode
{
// QuadTree
string UUID { get; }
IQuadTreeNode Parent { get; set; }
IQuadTreeNode Child1 { get; set; }
IQuadTreeNode Child2 { get; set; }
IQuadTreeNode Child3 { get; set; }
IQuadTreeNode Child4 { get; set; }
bool isLeaf { get; set; }
int Level { get; set; }
// Functions
void Split();
void Merge();
void Expand(int level);
int NumberOfLeafs();
IQuadTreeNode[] Leafs();
}

Additional to the “core” elements we store the very basic information about a plane in the node.

{
// QuadTree
string UUID { get; }
IQuadTreeNode Parent { get; set; }
IQuadTreeNode Child1 { get; set; }
IQuadTreeNode Child2 { get; set; }
IQuadTreeNode Child3 { get; set; }
IQuadTreeNode Child4 { get; set; }
bool isLeaf { get; set; }
int Level { get; set; }
// Plane
Vector3 OuterVector1 { get; set; }
Vector3 OuterVector2 { get; set; }
Vector3 OuterVector3 { get; set; }
Vector3 OuterVector4 { get; set; }
Vector3 CenterVector { get; }
float Width { get; }
float Azimuth { get; set; }
float Elevaltion { get; set; }
// <- Custom member that stores the vertice/triangle/… data
// Functions
void Split();
void Merge();
void Expand(int level);
int NumberOfLeafs();
IQuadTreeNode[] Leafs();
}

“OuterVector1-4” store the non-normalized outer edges of a plane/node within [-1|+1] range. Each vector (or vertice) lies on one of the planes of the non-normalized cube within its [-1|+1]. Normalization and radius are not involved yet.
“CenterVector” is the center of the plane. It is not stored but calculated on demand using OuterVector1 and OuterVector3.
“Width” is the distance between two OuterVectors which together connected make the edge of the plane, not the diameter. E.g. OuterVector1 and OuterVector2. Its also calculated, not stored.
“Azimuth” and “Elevation” are the rotations required to rotate the (normalized) plane so that its (normalized) CenterVector now lies on Y-up. Although these values could also be calculated on demand we store them once at creation of the quadtree node due to the computations required.

All further values, e.g. OuterVector1-4 values normalized or multiplied by the sphere’s radius can be derived. However it depends on the caching/implementation strategy if these additional values (normalized ones, final worldspace position) should be stored as well in the node or calculated on demand.

Plane Pipeline

In general when creatig a cubesphere we multiply all cube’s vertices on its surface in [-1|+1] value range by normalizing them first and then multiply them by the desired radius. Additional noise can either be added after normalization or after multiplying by the radius.
1.1 Cube plane in [-1|+1]
1.2 Normalize all vertices
1.3 Multiply vertices by radius
1.4 Multiply vertices with +/-noise value for terrain
or
2.1 Cube plane in [-1|+1]
2.2 Normalize all vertices
2.3 Multiply vertices with +/-noise value for terrain
2.4 Multiply vertices by radius

Approach 1.x might be prefered as chance is lower we get into precision issues and working with height information in the right radius scale might feel more intuitive.

But:

I am torn what an overall best strategie is. Maybe this also depends on the question if you are going to calculate your plane vertices on the CPU or GPU (and if on the GPU, how you use it to render the result).

GPU)
A previous approach here was to calculate the vertice worldspace position on the GPU using a compute shader. The calculations result remained on the GPUs memory and rendering these positions happended in the shader. In that case we needed vertices being prepared on the CPU only as “container” till that moment when they shall be rendered (and then displaced by the shader using the GPU data). In that case all planes can initially and throughout all planes consequently be created in e.g. Y-up. for all vertice data, as final positions are defined individually by GPU data.

CPU)
In this scenario you need, at some point, your vertices to be at their final worldspace position, including rotation and normalization of the plane and all vertices in the desired world radius.
This leads to the question if you really want to create your plane vertices initially in the Y-Up direction (and centered at e.g. 0,1,0) when you “know” that during the planepipeline process you have to rotate your plane (or all vertices) to one of the directions of the cube’s faces, center them to the node’s CenterVector, then normalize, apply radius and apply noise.

The process would read:
1.1.1 Create plane in Y-Up (centered at 0,1,0) (in already correct width so that it fits with depth of node)
1.1.2 Rotate planes/vertices to face into the same direction as the quadtrees centervector normal.
1.1.3 Center the plane to the CenterVector
1.2 Normalize all vertices
1.3 Multiply vertices by radius
1.4 Multiply vertices with +/-noise value for terrain
1.5 Calculate Elevation & Azimuth and store it for later use.

So we apply the normalization and terrain noise after the plane has reached its “desired” position on the cube (non-normalized) by using rotation and such
So the question is, why do we want to create the plane data in Y-up at (0,1,0) instead of simply creating the vertice positions directly in the right rotation and centered at centervector on the cubesphere’s [-1|+1] surface during initial creation of plane?
One argument would be that that at some point you want the plane in Y-up for AABB collision tests (e.g. for NavyFish’s LODSphere strategy). However Azimuth and Elevation can still be calculated by using just CenterVector.normalized. Am I overseeing something?

CPU vs GPU

As we desperately need some of the plane data on the CPU for LODing (at very least the bounds of the plane that holds the plane at its world position (including normalization, radius and noise) I feel I want this time to stick to CPU creation of the plane when it comes to vertice data. Knowing that calculating these planes take longer, my feeling is that having a good LOD in place with threaded CPU based plane creation might be more important. Terrain data taking longer to load worries me less than FPS spikes due to asynch readback of CPU data or non-sufficient LOD in place for very depth quadtree situation where the observer is close to the terrain surface. And I would like to to test how far I can push details down when using CPU’s double precision through most of the process to create the plane, as this up to now payed out a lot when re-implementing Unity stuff with hardly and performance impact (double vs. float on CPU).

So, how about the GPU? Creating material textures is obvious. But, I wonder if some clever approach could be found to use the GPU for more detail none the less.
I have to questions in my head and would wonder about your thoughts.

  • Do you know how likely it is to achive some identical noise implementations both on the CPU and GPU so that based on same parameters they deliver identical results? I know it feels odd to create the vertice position on the GPU and do the same to get to the normal map (which you want for higher details) on the GPU.
  • I wonder if there is a good way to use the GPU to interpolate rawer CPU data for finer details. So you provide the result of the CPU calculations (e.g. a plane in 12x12 resolution) to calculate the details between them. I am currently stuck at thinking you will need again to achive identical noise implementations because otherwise, even though it wouldnt be a problem to create details once, you would notice terrain changes when splitting vertices where GPU results wouldnt reflect CPU vertice positions when creating them for the next quadtree depth.
    But maybe you have some ideas or thoughts about this?

Looking forward to anyones comment or thought.
PS: Hope everyone is healthy and fine!!

1 Like

Anyone still working with Unity3D and procedural spherical terrain?
I would like to discuss approaches as I currently try to hunt down a bug that drives me crazy in the normalmapping.

The cubesphere looks ok when using simple vertex normals without normal mapping. No noise applied yet, all normals are simply the normals of the vertices positions.

When adding normal mapping and the code to calculate the normalmap of course, there is a lighting issue at the edges. Note that I added a skirt around the positional data for the normal calculation.
So when I for example apply a 32x32 map, the positional data is a grid of 34x34.
The code to calculate this data is the same as being used for the vertices.
So I calculate a flat plane between [-1|+1], then normalize the plane, then multiply by the radius.
This vector position data is used for vertice positions and for the positional data (so a finer grid of positions) for normalmapping. Except that for the normalmap, the skirt vectorswill be outside the [-1|+1] range (e.g. [-1.33|+1.33]) during the above mentioned calculation strategy. I am not sure if this is an issue, my understand is that this shouldnt be a problem after normalization, but I am not completely sure.
As the bug I search for only happens at the edges of the planes, so it might be indeed the skirt that causes the issue.

The normal map applied as texture where you can see the RGB normals get off…:

Anyone here faces this issue to to discuss the code to calculate the normals and the skirt topic?

1 Like

There are two things I can think of, first the positional data should be 2^n+1 not 2^n+2, the other issue is the orientation of the planes you are using in your cube map, the normal calculation needs to orientate the same way as the sides next to it. If you are calculating topLeft-topRight-bottomLeft for your normal and the side next to it doesn’t match up you get a seam.

1 Like

There’s also the issue of mipmapping and causing bleeding in your normal map, which this kind of looks like.

1 Like

EDIT 2020-06-20:
Looks like it has been indeed been bleeding and Japa was right. I fixed a bug in the uv code and afterwards the shadowing showed up on all edges, not like in previous screens. But it was still there, also when looking at the texture in the Unity editor while runnign the scene. However after saving the normalmap texture to disc and looking at it in photoshop I noticed the wrong edges were gone. So the Unity3D editor can really fool you here.

So I did a little bit of slight padding of the UV coordinates based on the textures resolution (I have to refine the values a little bit more these are just the initial testing values that delivered good results), and the issue was gone.

                // Calculate edge padding based on normalmap resolution to avoid texture bleeding
            float padding = 0;
            if (constants.nPixelsPerEdge > 0 && constants.nPixelsPerEdge <= 64) padding = 0.01f;
            else if (constants.nPixelsPerEdge > 64 && constants.nPixelsPerEdge <= 128) padding = 0.005f;
            else if (constants.nPixelsPerEdge > 128 && constants.nPixelsPerEdge <= 256) padding = 0.002f;
            else if (constants.nPixelsPerEdge > 256) padding = 0.001f;

            // Create uv 
            Vector2 uv = new Vector2();
            uv.x = Mathf.Clamp(col / (float)(constants.nVertsPerEdge - 1), padding, 1 - padding);
            uv.y = Mathf.Clamp(row / (float)(constants.nVertsPerEdge - 1), padding, 1 - padding);

3 Likes

I’ve finished the initial pipeline to render the spherical planes based on a quadtree, right now all in the main thread which is obviously only a temporary solution. Without LOD yet as I first want to focus on the plane pipeline.

I got to the point where I now have to decide fundamentally in which direction I want to go.
A) remain on the CPU and introduce multithreading by adding a worker thread to calculate the plane
B) stick to calculate planes on a main thread but calculate parts of the pipeline on the GPU.

This would be the options how to change the pipeline.

Option A)

  • Easy to introduce double precision
  • Data fully available on CPU
    o A seperate thread can be a pain
  • Probably always slower to calculate postional data based on noise

Option B)

  • Probably always massively faster than on CPU
    o Depending on level of GPU usage multiple asynchronous data transfers between CPU-GPU-CPU required
  • Impossible (?) to introduce double precision.

Even though Option A) seems to have more cons I feel torn to use the GPU. With the latest Unity version asynchronously getting the data back from GPU to the CPU seems have become stable (haven’t tried yet). However I would feed sad to let the double precision go.
If there would be a way that a ComputeShader can handle double precision vector data to calculate double precision noise that would be awesome. However I think this is still not given right?

What direction would you move to?

1 Like

Pretty sure compute shaders are capable of double precision calculations, you just lose on performance.

1 Like

As far as I’m aware you still can’t use double precision on consumer GPU’s, double precision is only enabled on workstation and datacenter GPU’s like Quadro, Tesla, and AMD’s equivalents.

1 Like

More specifically, or at least last time I checked, only some simple instructions ( basically construction and additions ) are supported in double precision on consumer GPUs.

1 Like

For reference

https://en.wikipedia.org/wiki/Feature_levels_in_Direct3D#Support_matrix
https://docs.microsoft.com/en-us/windows/win32/api/d3d11/ns-d3d11-d3d11_feature_data_doubles
https://docs.microsoft.com/en-us/windows-hardware/drivers/display/directx-feature-improvements-in-windows-8#dblshader

1 Like

Thanks for the additions Keith, Flavien and cybercritic!! :+1:
I decided to go the GPU route as I really cannot let the parallelism of the GPU asside for the terrain calculation. Even though I might not be able to apply double precision completely on the GPU / ComputeShader, the more I think about it the less I am sure if “just” having single-precision floats would be really an unsolvable problem when using it for the meshes of a cubesphere where terrain is splitted by a quadtree mechanism. For the individual meshes it shouldnt be a problen as they will either need to cover a large range at low precision or a small scale with high precision (where float should be OK). Only thing that bothers me a bit is that typically the noise algorithm is fed with the positional vectors after normalizing them and applying the spheres radius. Where then some precision issues could occur. But maybe there is some clever workaround for this. To be seen.

So I started to implementing the ComputeShader to create the position of the vertices now.
My original idea which I also posted in the above picture was to create the normalized vertice positions on the CPU and dispatch it to the GPU to calculate the noise values, however probably it is way more efficient to also directly calculate the normalized positions on the GPU. This would reduce the data transfered from CPU to GPU massively which I understand is still a botteneck.
So my rough idea is to implement the kernels like that.

Set PatchConstants, BodyConstants and OutputBuffer to ComputeShader and Dispatch to GPU
kernel[0] = Calculate normalized positions and apply noise (vertices-resolution)
AsyncGPUReadback vertice positions from CumputoBuffer (GPU) to Array (CPU)

Set PatchConstants, BodyConstants and OutputBuffer to ComputeShader and Dispatch to GPU
kernel[0] = Calculate normalized positions and apply noise (pixels-resolution with skirt)
kernel[1] = Calculate normals/normalmap (pixels-resolution)
kernel[2] = Calculate slope (pixels-resolution)
kernel[3] = Calculate material texture (pixels-resolution)
AsyncGPUReadback normalmap and texturemap from ComputeBuffer (GPU) to Array (CPU)

Tasks of kernels 0 and 1 might be combined but currently I am thinking to split them to keep some structure in the shader, however this slows down too much I might combine then.

I also decided to for the moment calculate vertice positions (probably 16x16) and pixelpositions (256x256) for later normalmapping in two separate steps instead of reusing positions.

Once I have done the vertice position I hopefully can get some feedback here how well reading back the data asynchronous from GPU to CPU works in Unity with the latest version. At least the first test-implementation works well.

EDIT [2020-07-20]
I implemented the first ComputeShader kernel which calculates the vertice postions and applies noise on the GPU.
The ComputerBuffer with the vertice positions is asynchronously read back from GPU to CPU and the mesh is created. Looks promising as asynchronously fetching back a computebuffer seems to work very well. I need to add the normals to the vertice data for which I need to create a separate buffer for this which seems like an overhead (as they are just spherical normals) but then the vertices part would be complete.

Nextup then are additional kernels to calculate the higher resolution normalmap. It will be interesting so see if fetching multiple buffers after multiple kernel steps back to CPU remains that easy.
I hope this is going to add some eye-candy, so far only a boring mesh :-).

EDIT [2020-07-21]
I added the normals to the outputbuffer. Luckily it is not necessary to create a separate buffer for each value. The documentation lacks a bit in how to do this, and I saw a few tutorials which created a separate buffer for each value to read it back into an array. However using a struct in a similar way like when you dispatch your stuff to the CPU simplifies things.

This is how you do it when you want to read a buffer which constists of vertice positions and their normals

In your computeshader:

// The structure of the output computebuffers
struct OutputStruct
{
float3 position;
float3 normal;
};

In C# you create the same struct:

struct OutputBuffer
{
// Vertices data
public Vector3 position;
public Vector3 normal;
}

Creating the buffer before dispatching:

patchData.outputBuffer = new ComputeBuffer(patchConstants.nVerts, 12+12, ComputeBufferType.Default);
// Output buffer contains
// position (float3 = 12 bytes)
// normal (float3 = 12 bytes)

And when reading the buffer when you are done on the GPU and checked that the buffer is ready to be read, you create an array of these structs and read out both values and assign them to their proper target array.

// Read back data
OutputBuffer[] outputBuffer = AsyncGPUReadbackRequest.request.GetData(0).ToArray();
for (int i=0; i <outputBuffer.Length;i++)
{
patchData.vertices[i] = outputBuffer[i].position;
patchData.normals[i] = outputBuffer[i].normal;
}

Not sure if there is a little room for improvement to avoid creating a separate array to iterate on, however it does the job and is pretty simple.

1 Like

The implementation of the GPU pipeline to render a plane is becoming slowly complete for a first iteration

CPU prepares some buffer constants and calculates the UV array and the triangle indices.
GPU Kernel 1 processes the vertice positions (incl. noise)
GPU Kernel 2 processes the positions (incl. noise) for the normalmap
GPU Kernel 3 processes the normalmap, the slope of the terrain and the surface texture.

The data is dispatched from CPU to GPU and read back asynchronously twice.
First dispatch is to GPU kernel 1 and asynchronously read back. After that
second dispatch is to GPU kernel 2 and asynchroundly read back from GPU kernel 3 (which works on the result of kernel 3)

The splitting and merging of the quadtree nodes and the planes already works too, so I can now move on to do the splitting and merging based on the camera position and some culling mechanisms (which were already discussed here, e.g. @NavyFish 's HorizonCulling mechanism.
This will then be the moment to see how well Unity3D’s asynchronous read mechanism from GPU to CPU really performs when continously rendering of new planes is going on. But so far it looks promising.

PS: There was a small but silly bug in previous implementations. For documentation purpose (as parts of the code were posted here previously):
When working with a buffer in the GPU which is meant to hold positional data including a skirt (to calculate the normal map later) it is likely, depending on how we calculate the positional data, id.x = 0 and id.y=0, does not correspond in the first skirt position but the first position of the plane. Therefor if we work on data that holds a skirt, when using the algorithms that were previously posted here by me or other we have to shift the position by one spacing.

Therefor I invented a function parameter that tells the function if it the data being worked in includes a skirt, and shift the positional data then in that case.

// First calculate the 'cube space' coordinates of the vertex:          
float eastValue = x - ((nPerEdge - 1) / 2.0);                        // eastValue now ranges from -.5*nVertsPerSide to +.5*nVertsPerSide
eastValue *= spacing;                                                // eastValue now ranges from -.5*patchWidth to +.5*patchWidth (patchWidth is not actually a defined variable)
if (hasSkirt) eastValue -= spacing;                                  // eastValue is shifted back into right range if indexes include skirt positions so that the skirt starts outside -.5
float3 cubeCoordEast = eastDirection * eastValue;
// Do the same for the "north" direction:
float northValue = y - ((nPerEdge - 1) / 2.0);
northValue *= spacing;
if (hasSkirt) northValue -= spacing;
float3 cubeCoordNorth = northDirection * northValue;

I got further than you @joergzdarsky, but honestly the best advice I can give you is to abandon this, you can make multiple games in the same time-frame as this experiment. The key is to use stock technology and make games with themes, topics and gameplay that you want to make.

Two years from now you might be implementing ray tracing and even if you get the tech done in this cycle, you are looking at these planets being an object that you could never hope to populate, so years taken to have a pretty and realistic backdrop, to a game that could have had the same gameplay as a 2d game.