Procedural Terrain Rendering How-To

I know what you mean @cybercritic, working on these topics is pretty time consuming. In fact especially when being one person working on such stuff, its an endless task literally. I am aware of this since (checking the archive) first iterations started somewhere near Unity3D 5.3, so its been awhile. And I would agree if my goal would be to make a game out of that.
However this is not my intention, which goes more towards prototyping, testing different implementation strategies, or how now Unity3D can be applied to such scenarios.
And discuss that with others. I went to the asset store to get some cockpit asset (as I lack the capability making 3D models which you sometimes want here or there) for testing purpose.
I’ve been thinking about getting some space station prefabs too, however I find the idea more interesting to create them procedurally too. But thats it with assets basically.

And last but not least, don’t forget that the procedural planets are just part (although they can be a huge topic themselfs) of what I work on.
Content and idea where I want to go to of my prototype is basically a sandbox environment similar to Space Engine (however very very far (!) from its visuals and quality), testing for example how to manage objects on such large scale.
I’ve been working on that topic (managing objects, large scales and velocity) for the last 1-2 years and finally got to an approach within Unity I am pretty satisfied with and with I am most likely will stick to now. Navigating in it basically feels a little similar to Space Engine which is what I wanted to achive, therefor I am now going to enrich it with procedural content where I currently use only debugging objects, and especially want to find out how well planetary objects based on some quadtree/cubesphere stuff works combined with what I did so far. Therefor I depend on having the control on for example how planets are being created, of how many meshes or gameobjects the surface consists etc, as there are different ways how I potentially can integrate them to move them trough space.

Basically its for me about getting the math done, into how to built an engine while working out how Unity3D works underlying looking at decompiled classes and API documentation, testing what you have to workaround with todays limitations, and applying this to a space scenario. Or simply: Continuous learning :slight_smile:
While doing this I met Chris Thorne who’s been the originator of the Floating Origin approach, and we have been discussing and sharing new ideas and some code on a gitlab repo, and I had the chance to look onto some of its papers (he is a nice guy!) and its great fun.

So, using only predefined stock stuff out of some asset store and taking the shortcut would not be right way to go. Infact it would be pretty boring :slight_smile:

2 Likes

@NavyFish might recognize this… :wink:

1 Like

I finished implementing the initial pipeline where the planetary terrain is rendered on GPU and read back asynchronously back to CPU, organized in six quadtrees, using the cubesphere mechanism to create the spherical terrain, and which uses the LODSphere strategy @NavyFish came up with for LOD (=split or merge).

Missing is some optimization, e.g. currently I only calculate only one terrain at once until the next terrain segement is calculated for debugging purpose of the quadtree parsing, however queues are already implemented. I just pick out only one terrain calculation task out of the queue at once to put it into the calculation pipeline instead of more in parallel. Also missing is HorizonCulling currently and lots of other things of course. But overall I am satisfied with the initial non-optimized performance using the AsyncGPUReadback API.

What I now stumble across is the floatng point precision issue which becomes visible rougly at node level 16-17 for a planet at earth size (6371km radius) as in this video. Terrain patches are 64x64 vertices and the material texture and normalmap texture being 256x256.

The pipeline is roughly

  1. Sent PatchConstants data (basically the quadtree node information, vertice-spacing etc) and BodyConstants data (e.g. planetary radius, noise data) from CPU to GPU.
  2. Calculate Vertex positions (GPU)
    2.1 Calculate position in [-1|+1] cubespace
    2.2 Normalize position in [-1|+1] spherical space
    2.3 Apply radius to bring sphere to “realworld” size
    2.4 Calculate noise (by using realworld size spherical positions) and add/substract some height.
  3. Calculate “Pixel-Positions” (for later normalmap texture and material texture) (GPU)
    3.x Same process as for the vertex positions
  4. Calculate normalmap texture out of “Pixel-Positions” (GPU)
  5. Calculate material texture. (GPU)
  6. Asychronous read data and texture from GPU to CPU and render the terrain.

I wonder where to look first to eliminate the floating point precision issue and how huge the problem overall is. My first thought is that it might occour very early when working with floats being the edge coordinates in the quadtree node, and by then calculating the flat terrain in [-1|+1] cubespace.
My impression is that the decimal positions already get insufficient here which then becomes more visible when this precision issue is projected to real-world space applying the radius, and using this data as noise input. Something that might be visible be debugging the node data for a deep-level node (e.g. level 17).
If this would be the case then doing some tricks in the ComputeShader like shifting terrain back to the origin etc. might not help when the problem is already earlier in the process.

I wonder if it would be worth to try out to change the quadtree node data from float to double, send this information to the GPU, calculate also the [-1|+1] cubespace terrain in double on the GPU, also the normalized spherical terrain in double, and also the “real world” position on the GPU in double before then converting it back to float to feed the noise algorithms. Up to that point all math applied to doubles on the GPU would as far as I currently oversee it only float and float3 addition/substraction/division/multiplication and normalization, which hopefully the GPU would be capable of.

Any thoughts? Am I on the wrong track and this would be an impossible task or the problem being much bigger requiring a totally different strategy?

1 Like

That is mostly likely your culprit. Also, doing on the GPU is gonna be tricky. Unless you’re targetting professional cards only ? Because consumer ones only support multiply-add in doubles. And to calculate your noise algorithm’s input position ( terrain coordinates in place space ), you’re probably doing a normalize() at some point. Which involves a recriprocal square root, and IIRC that’s not supported in doubles in hardware ( basically it’ll do the operation in fp32 and you won’t see any improved precision ). But it was many years ago when I checked that stuff, so it may have changed on newer GPUs since then. Might be worth a retry.

2 Likes

Thank you very much for your reply @INovaeFlavien / Flavien .
I’ve been checking google and my short research resulted in the same (only limited math in combination with doubles) like you mentioned. And even if possible I do not really want to restrict my implementation only to most current GPUs. However eventually I will try to add some simple double values and do some basic math into the ComputeShader to see at least if the compiler complains or not and eventually try to debug some results.

May I ask if you utilize the GPU to calculate your planetary patches, or at least for parts of the process? Not asking for details, just a simple yes or no (probably yes) so that I can have at least a little hope that it might make sense to investigate into that further more on my own.

tl/dr: Initial thoughts about the float precision issue and maybe how to overcome it.

So I set a breakpooint at my patch creation for a level 18 node to check the data of the node.

node.Level = 18
node.OuterVector1 = (-0.1113586,1,-0.4103546)
node.OuterVector2 = (-0.111351, 1,-0.4103546)
node.OuterVector3 = (-0.1113586,1,-0.4103622)
node.OuterVector4 = (-0.111351, 1,-0.4103622)
node.CenterVector = (-0.1113548,1,-0.4103584)
node.Diameter = 1.078959E-05

patchConstants.spacingPixels = 2.99192E-08
patchConstants.spacingVerts = 1.211015E-07
patchConstants.nPixelsPerEdgeWithSkirt = 258
patchConstants.nVertsPerEdge = 64

This means that the math to create the vertice and pixel positions initally in [-1|+1] cube space
on GPU (before normalizing and pushing the data out of [-1|+1] to real world radius) has to handle the following situation (which is currently happening on GPU in float3):

OuterVector1(x) -0,1113586
OuterVector2(x) -0,111351
Distance (x) -7,6E-06
Split by 256 -2,96875E-08

Indeed this looks like being dangerously close to float precision issues already here at the very early point within my node data.

As my approach of using a cubesphere typically created in a [-1|+1] cube first, then normalized and applying radius, and organizing this in a quadtree, I made a table to check how precise I am currently with that.

Looks surprisingly not toooo bad for a test implementation just looking at the resolution, but not optimal, especially when you think about rendering bigger planets.

So where am I, asuming we won’t be able to change the full GPU workflow in the ComputeShader to double precision?

The process currents consists of:

0.0 Check/Update Quadtree Node / Create patch Parameters | Resource-Impact: LOW
1.0 UV-Coordinate and Triangle-Incides arrays | Resource-Impact: LOW
2.1 Initial Vertex/Pixel Position arrays [-1|+1] - flat cubespace | Resource-Impact: LOW
2.2 Updated Vertex/Pixel Position arrays [-1|+1] - normalized | Resource-Impact: LOW
2.3 Updated Vertex/Pixel Position arrays [realworld] - multiplied by radius | Resource-Impact: LOW
3.0 Noise array (SimplexNoise using Vertex/Pixel Position arrays[realworld]) | Resource-Impact: HIGH
4.0 Updated Vertex/Pixel Position arrays [realworld] - add/subsctract by noise | Resource-Impact: LOW
5.0 Normals and NormalMap-Texture (by calculating normals out of Vertex/Pixel Position arrays [realworld] input) | Resource-Impact: MEDIUM-HIGH
6.0 Material-Texture (using noise data and normals as input) | Resource-Impact: MEDIUM

with:

0.0 Check/Update Quadtree Node / Create patch Parameters (float) (CPU)
CPU -> GPU
1.0 UV-Coordinate and Triangle-Incides arrays (float) (CPU)
2.1 Initial Vertex/Pixel Position arrays [-1|+1] - flat cubespace (float) (GPU)
2.2 Updated Vertex/Pixel Position arrays [-1|+1] - normalized (float) (GPU)
2.3 Updated Vertex/Pixel Position arrays [realworld] - multiplied by radius (float) (GPU)
3.0 Noise array (Precision: float) (GPU)
4.0 Updated Vertex/Pixel Position arrays [realworld] - add/subsctract by noise (float) (GPU)
5.0 Normals and NormalMap-Texture (float) (GPU)
6.0 Material-Texture (float) (GPU)
GPU -> CPU (asynch)

As the precision issue seems to be originating in 0.0 within the node data being floats, and asuming the GPU can only handle float, my initial idea would be to change the process by moving certain parts back to the CPU and change it to double. Basically the CPU would calculate the round spherical surface and send it to the GPU to calculate noise (step 2.3->3.0) and normals.
At this point where send positional data from CPU to GPU I would have to downgrade double back to float again for the ComputeShader (but probably keep my double precision data on the GPU).
The GPU would use this downgraded float as input data to calculate noise and furthermore the normals and normalmaptexture.

0.0 Check/Update Quadtree Node / Create patch Parameters (double) (CPU)
1.0 UV-Coordinate and Triangle-Incides arrays (double) (CPU)
2.1 Initial Vertex/Pixel Position arrays [-1|+1] - flat cubespace (double) (CPU)
2.2 Updated Vertex/Pixel Position arrays [-1|+1] - normalized (double) (CPU)
2.3 Updated Vertex/Pixel Position arrays [realworld] - multiplied by radius (double) (CPU)
CPU -> GPU
3.0 Noise array (float) (GPU)
4.0 Updated Vertex/Pixel Position arrays [realworld] - add/subsctract by noise (float) (GPU)
5.0 Normals and NormalMap-Texture (float) (GPU)
6.0 Material-Texture (float) (GPU)
GPU -> CPU (asynch)

The interesting question is, will there be a good way to cast the double back to float on the CPU before dispatching it to the GPU so that it can act as valid input for the noise algorithm there. And furthermore, would the downgraded spherical position data be sufficient, after applying the noise result to them on the GPU to continue to work with that data (downgraded position data and noise result)
to calculate the normals, without reading the noise back to the CPU inbetween.

Maybe one solution would be to multiply the spherical [-1|+1] positions (in double) not by the radius but by some value to shift numbers and that prevents the loss of digits when casting to float, and send these temporary positions to the GPU’s noise function to work with that instead of the realworld ones (that become might become too big for a float when getting closer to the surface).

Another solution might be to create two temporary position arrays sent to the GPU, one HIGH one that holds the double data before the digits (so xxxxxxx.0) and one LOW one that holds the data after the digits (so 0.xxxxxx), call the GPU’s noise function twice with both variants, and stack the result.

Curious about anyones thoughs, has someone of you thought about preventing loss of precision in a planetary scenario already?

That won’t work, but scaling the values you’re also scaling the threshold at which you have the precision issues, so in the end it doesn’t matter what the scale is: FP32 usually only offers 6 digits of good precision, aka. you’ll get the same issue whether it’s at 123456.0 or 0.123456

I can’t really explain how we’ve done it in the I-Novae engine, but I can say that you’re on the right tracks. In particular the part about this might be worth more research / investigations:

We handle all the noise / elevation generation on the GPU, it’d simply be too slow for CPUs.

2 Likes

Thank you very much Flavien for the answer. It makes me optimistic to continue.
I am going to give it a try to rework the initial function to double precision and start working on feeding the noise function with two different HIGH and LOW values and investigate on that. Really curious how this will work.

1 Like

Sometimes procedural stuff is really fun, you always stumble across unexpected outcomes.
I started to implement to split the positional coordinates (created on the CPU) of each plane position which feed the noises xyz parameters (on the GPU) into two, in order switch on the CPU to double precision and upload this in splitted positional data to the GPU and call the noise function there twice, and then stack the result.
So what previous has been one vector3 position array (which is later uploaded in an buffer to the GPU) now becomes two, with one array storing the numbers before the decimal,
and one the numbers after the decimal).

Vector3[] positions

became

Vector3[] positions_HIGH
Vector3[] positions_LOW

So a previous position that was e.g. (142.1847, 1432.298,76.73849) now becomes
in position_HIGH (142, 1432,76)
in position_LOW (0.1847, 0.298,0.73849)

For testing purpose I first checked each noise result for position_HIGH and position_LOW separately before trying to stack them and stumbled across something unexpected here.
position_HIGH rendered just fine and didnt look much different to the previous unsplitted result. Which was more or less expcted as I just loose a little bit of decimal precision however the numbers are more or less the same, and at high scale this does not matter.

Then I rendered to position_LOW data and expected a totally randomly looking fine-grained result, as the data after the decimal are overall quasi-random when you discard the data before the decimal,
and all very small when used at a higher distance.

However in the result there was an unexpected pattern. Basically the plane rendered as expected (fine grained randomly looking plane) however some… bumps… appeared.
They do not keep their position but which is expected, remembering I just work with the data after the decimal), but they appear on all distance levels.

What you can see is that at certain positions the noise results tend to become suddenly very high I guess. I asume this is some sudden high noise output at some points for certain reasons. The pattern around these (white/black spot) eventually a result of the directional light and shadow.

The way I split one float into two is simple.

/// <summary>
/// Removes the decimal positions within the vector.
/// </summary>
/// <param name="vector"></param>
/// <returns></returns>
private static Vector3 ToHigh(Vector3 vector)
{
    Vector3 result = new Vector3();
    result.x = (float)Math.Truncate(vector.x);
    result.y = (float)Math.Truncate(vector.y);
    result.z = (float)Math.Truncate(vector.z);
    return result;
}

/// <summary>
/// Keeps just the decimals of the vector and sets the value before the decimal to zero.
/// </summary>
/// <param name="vector"></param>
/// <returns></returns>
private static Vector3 ToLow(Vector3 vector)
{
    Vector3 result = new Vector3();
    result.x = vector.x - (float)Math.Truncate(vector.x);
    result.y = vector.y - (float)Math.Truncate(vector.y);
    result.z = vector.z - (float)Math.Truncate(vector.z);
    return result;
}

I have not yet moved to change the vector data to double precision as I feel Ifirst want to know here this comes from. Even though I have hope that this issue might disappear when moving from float to double, as the positional data in float is generally relatively large so that often the decimal position_LOW values are close to zero as there is not enough precision for large numbers in float.
Often the float values of a vector require so much space so that the decimals are completely zero which then would lead to (0,0,0) as input data in positions_LOW (I guess perlin noise has a certain behavior for (0,0,0) as input parameter). Maybe that is the already the reason.

Howevery its a funny observation inbetween to hunt for the cause :slight_smile:

It will be in interesting challenge how to stack both noise results of position_high and position_LOW later again. Basically I’d say I have to just add both noise results and divide by two, but weight the high noise value way more than the low one (maybe with just 1%). So something like

noise = (noise1 * 0.99f + noise2 * 0.01f ) / 2;

However not sure if that is to simple… to be seen.

4 Likes

Im back :slight_smile: How is it going guys?

4 Likes

Here are a short Video.

It getting better and better, but for the terrain i struggle, i dont love noise functions :smiley:

2 Likes

@montify maybe this can help.

1 Like

As I am a little afraid that this post might get lost at some point (who knows what changes to the forum might occur) with all the valuable information in this thread, I am currently trying to export it to PDF or something, however currently this fails. Has anybody of you been able to export this thread somehow or might know a trick how to do it?

PS: Have a happy new year 2022 everyone! Take care and stay safe and healthy!!

2 Likes

I afraid it’s only possible with step-by-step copy paste, because of how this forum works and loads up next-posts after a certain limit of visible posts.

You, guys HNE too!

Cheers!

1 Like

I try to continue the planet journey,is anybody active ,it’s been a long time

1 Like

Hello,

here is my nth Iteration of my Planet Renderer.

Now i implemented some sort of Collision Detection as you can see in the Video.

And i switched from generating BoundingBoxes to use the 4 points of the Terrain Patch to generate a Plane, it Culls the same amount of Nodes, but i can throw away the BoundingBox code, because of course, i have the Terrain Patch already.

1 Like

I have, yes. Shoot me a DM if you’re still looking for an archival copy.

Also, hello :slight_smile: