Procedural Terrain Rendering How-To

Guys, looks like you ignore me :slightly_smiling:
Can someone explain how do you get working normal mapping? Step by step. Please.
ATM i calculate four direction of NM to get good slope…

float l = patchPreOutputSub[(id.x + 0) + (id.y + 1) * nVertsPerSideWithBorderSub].noise * constants.lodLevel;
float r = patchPreOutputSub[(id.x + 2) + (id.y + 1) * nVertsPerSideWithBorderSub].noise * constants.lodLevel;
float u = patchPreOutputSub[(id.x + 1) + (id.y + 0) * nVertsPerSideWithBorderSub].noise * constants.lodLevel;
float d = patchPreOutputSub[(id.x + 1) + (id.y + 2) * nVertsPerSideWithBorderSub].noise * constants.lodLevel;
…
float xdelta = ((l - r) + 1.0) * 0.5;
float ydelta = ((u - d) + 1.0) * 0.5;
float zdelta = ((r - l) + 1.0) * 0.5;
float wdelta = ((d - u) + 1.0) * 0.5;
…
float3 xnormal = float3(xdelta, ydelta, 1.0);
float3 ynormal = float3(ydelta, xdelta, 1.0);
float3 znormal = float3(zdelta, wdelta, 1.0);
float3 wnormal = float3(wdelta, zdelta, 1.0);
…
float xslope = 0.5 / max(dot(xnormal, float3(0.0, 1.0, 0.0)), 0.001);
float yslope = 0.5 / max(dot(ynormal, float3(0.0, 1.0, 0.0)), 0.001);
float zslope = 0.5 / max(dot(znormal, float3(0.0, 1.0, 0.0)), 0.001);
float wslope = 0.5 / max(dot(wnormal, float3(0.0, 1.0, 0.0)), 0.001);
…
float slope = min(min(xslope, yslope), min(zslope, wslope));
…
float4 color = ColorFunction(cpos.xyz, noise, slope);
float4 normalColor = float4(normal, slope);
float4 heightColor = float4(color.xyz, noise);
…

So results is:

  1. Normal map texture. Normals(RGB), Slope(A)
    RGB
    A

  2. Height map texture. Color(RGB), Height(Noise)(A)
    RGB
    A

Next step for me is surface shader…:
Normalmap direction artifacts
Colormp artifacts

While working on surface transition to start working on surface detail, I noticed that the precision problem again got me.
Planet sphere radius is 6371000 (earth size), plane mesh is 32x32, surface texture and normalmap is 128x128, maximum node depth is 18, maximum noise octaves 23.

At the highest nodelevel the textures begin to have vertical and horizontal artefacts (as well as errors in the noise calulcation), and positional jitter occours.

Guess the texture issues are due to float precision in the compute shader.
And the positional jitter could be due to the fact that each plane’s reference position (the sphere’s gameobject center) is at relative 6371000 distance. E.g. if I am directly before the plane surface, the referenced center coordinate (because the plane has no own gameobject) is at (0,0,6371000). It could be either that, or its again a precision problem in the shader, that time in the vertex shader, where the plane’s vertices are displaced from (0,0,0) reference to the planes worldspace center.

I really wished Unity would support double precision (and somewhen shaders too sigh. This would make live so much easier…

This is the exact same problem that I’ve been trying to tackle. The problem is that we run out of FP32 precision on the GPU. We could probably squeeze more precision if we start with a larger number on the first LOD level, but then that means the noise frequency will be too dense, or at least for me with the Simplex noise algorithm. So it’s not a solution.

There’s one method that has been suggested before, which is emulating FP64 on the shader, and it should be adequate to just pass emulated FP64 number to the noise function as an input and keep the rest of the fractal calculations as FP32 in order to get 3-4 more LOD levels without artifacts. I tried this myself with no luck, I probably did something wrong on the way. I will probably try this again soon.

I’ve also thought of another method. Why is the noise frequency so dense with large numbers? Could we maybe change the noise algorithm so that the noise has less frequency for larger numbers? Because then we could probably squeeze out more precision. Because we’re only utilizing part of the FP32 range.

There’s an interesting thread on the Outerra forums where cameni is talking about using integer noise function to get more precision. He’s sounds surprised why no one is doing this.
http://forum.outerra.com/index.php?topic=245.0

I came up with the code listed above only after everybody presented the same direct function for computing the noise and complained about the precision. So I wondered why people aren’t using an integer version that can use more bits for computation. However, my version uses 2D coordinates as the input, I can see it would be simpler for people to use 3D there. Well, it should work as you have described.

The code he’s talking about:

Suppose you have unsigned integer coordinates ranging from 0 to 2³²-1, used to address the location on the face.
The following code is for 1D, using a 1D pregenerated noise texture, but can be easily extended to more dimensions.

float interpolated_noise( uint x, uint level )
{
uint xm = x & (0xffffffffU >> level); //mask out the bits that would cause repeat in tex lookup
float c = ldexp(float(xm), level-32); //convert to 0…1 floating point value

//use c in noise tex lookup
float v = texture(noise1D, c).x;
//could be scaled here as well
return ldexp(v, -level);

}

Then you just sum up the values for how many levels you need to get the result.

The guy who asked the question replied back asking if this implementation is correct:

would you call that function on each axis of your face coordinate?
when you talk about face coordinates i’m getting the impression you mean a 2D grid coordinate. Is that not the case?
is this how you see it being used?

vec3f unitVector; // in range of -1.0 to +1.0
vec3ui intVector = convertFloatToUINT(unitVector);
// intVector now in range 0 to 0xFFFFFFFF
float n = 0.0;
for (int o=0; o<octaves; ++o) {
n += interpolated_noise(intVector.x, o);
n += interpolated_noise(intVector.y, o);
n += interpolated_noise(intVector.z, o);
}

And to make long story short, he says this implementation should pretty much work.

Now, this code is really confusing for me, especially the float to uint part, I have no clue how that works. But I hope my reply gives you some ideas on what you can try out.

I would love to hear from someone who managed to solve this problem!

I would also like to add that I solved the vertex precision error (where the mesh is bouncing and jittering around) by having 64 bit precision on the CPU and then when I upload the vertices to the GPU I subtract the camera position from the mesh, so that you get maximum precision for everything that is close to the camera. So basically the camera is always at origin (0, 0, 0) and I just shift everything around the camera for maximum precision.

Hi raRaRa,

thanks for the response. Will look into your first hints in more detail, it could take a while. The referenced thread is a nice read!! It gives some ideas what could help here.

Well I do something similar, but only when it comes to all positions on the CPU. The camera remains centered (well, say it is set back after a certain small distance) at (0,0,0) using floating origin. Each object is placed and moved around the camera so that the highest precision remains around the camera.

My planet has only one positional coordinate, the center. When I want to move my planet itself, I only move the center coordinate. Now, the planes are relatively positioned around this planet’s center coordinate (they dont have its own gameobject), but calculated within their own coordinate system on the shader, around (0,0,0). What happens is, the compute shader calculates the plane and stores all vertex positions around (0,0,0) center. Then the vertex shader takes the vertex positions and displaces them with the planes world space center coordinate. So for example the top plane’s vertices are displaced by (0,6371000.0f, 0) in the vertex shader, which could be the problem as you stated.

Indeed the only solution I can think of right now is to displace the vertices relatively to the camera, not to the planets center. But then I would need to have to ā€œcutā€ the planes reference to the planet’s gameobject, and each plane would need its own one, I think. Need to think about that too for a while, but yes considering the camera somehow seems like the best approach right now to me.

In my case, each patch has its own patch space, the center of the patch is always (0, 0, 0). Then I just move that patch to it’s correct place. I send the center coordinate of the patch in planet space to the shader and add it to the vertices in the vertex shader. That seems to be enough to get rid of the jittering.

BTW, I’ve created a Shadertoy project to experiment with the normal map artifacts. Feel free to play around with it. I’m planning to try emulating FP64 there and see if it helps to fix the artifacts.

The Shadertoy has one render buffer, which is basically rendering the height map to a texture, and then I use sobel fiter on the main shader to render the normal map. Click on the image and drag the mouse to the left and right to zoom.

https://www.shadertoy.com/view/4dKGRV

I found a forum thread today that discusses the normal map precision problem we’re experiencing, and it seems like he found one solution (that I kinda dislike). It seems like he had to generate the height map twice, one for the high part of the emulated double, and another for the low part of the emulated double. Otherwise the height map wouldn’t have enough precision to get rid of the artifacts in the normal map generation.

He tried exactly what I tried first, by just passing the emulated double number to the noise input, but that wasn’t enough. I can think of one case where it might be adequate, and that is if we don’t rely on the height map and just generate the normal map by calling the noise function, but it might introduce performance problem, since you would have to call the noise function for the surrounding texels. (left, above, right, below, etc.).

It would be awesome if @INovaeFlavien could share what method he used! :slightly_smiling:

I’m not sure it works. The problem is where this ā€œunitVectorā€ is coming from. If it is, as I suspect, the normalized position in planet space, then you’re already losing your precision there. It doesn’t matter if you do any nice tricks with integers afterwards. For it to work IMO you’d need to generate that unit vector in full precision.

Well no, I can’t really discuss implementation details, as it’s part of our core technology.

All I can say is that we support single floats, doubles or double-emulation. Depending on the complexity of the planet ( in subdivision depth ) and gpu, we choose one or another.

Thanks for the answer, and I completely understand.

Can you maybe give me a blinky if I’m thinking this correctly below? :slightly_smiling:

I have no clue why SkavenPlanet is sending the noise input coordinates to the normal map generator, unless he’s creating some kind of planet space normal map. Judging from his post, this should work if I send emulated double precision coordinates to the noise function, and also emulate the output, and then I can simply generate a height map that simply has the highHeight, lowHeight. I wouldn’t need to generate two textures since I’m using tangent normal map, so it shouldn’t require any unnecessary data such as the patch coordinates.

In my mind this should do the trick. But It’s always motivating to get verified by the professionals.

Thanks!

I don’t think he’s sending noise input coordinates to the normal map generator, although I could be wrong since I know nothing about how he implemented it. The normal map generator only uses the height values from the previous step ( procedural generation ) which itself uses the HP input coordinates. What he’s saying is that you need to retain full precision between these 2 steps. You use a temporary texture to store the height values, so one way to store height values as doubles is to split each value into two FP32 values, which then requires two textures, which gets recombined in the normal map shader to get the original high-precision value back.

Edit: sorry, reading back again, he’s indeed passing the input coordinates to the normal map generator; I have no idea why.

1 Like

But there shouldn’t be any need for two textures to store the double, because I could simply use the R and G components of the texture, for the low and high part of the emulated FP64 height value, right?

Thanks again!

Logically, yes. You can even use an RG texture to avoid wasting space. He used two textures because for some reason he’s also passing the XYZ values for the input coordinates.

1 Like

While looking into optimization options for the precision issue of noise and jittering and the hints of @raRaRa, I also came across the code that gave me headache for performance.

I decided to move to astroid rendering inbetween until a solution for the precision, but now on GPU (my previous asteroids were built on CPU). I rebuilt my planet code slightly so that it can render a planet or an astroid just by changing input parameters (enabling or disabling separate quadtree threading, separate LOD parameters etc), and gave it a try.

Noticeable was the drop in performance when adding around +100 to +200 asteroids into the scene. I profiled the issue and again it ended in the void OnRenderObject() function with the included DrawMesh call. The same method that was shown in the profiler when rendering a planet at close surface, consuming much CPU resources.
Commenting the Graphics.DrawMesh out I was at nice 60 fps (with of course nothing rendered).

void OnRenderObject()
{
    int QuadtreeTerrainRenderQueueCount = this.QuadtreeTerrainRenderQueue.Count;
    int layer = LayerMask.NameToLayer(GlobalVariablesManager.Instance.layerLocalSpaceName);
        // Render
        for (int i = 0; i < QuadtreeTerrainRenderQueueCount; i++)
        {
            // Draw Mesh
            Graphics.DrawMesh(this.QuadtreeTerrainRenderQueue[i].prototypeMesh, transform.localToWorldMatrix, this.QuadtreeTerrainRenderQueue[i].material, layer, null, 0, null, true, true);
        }
}

By using quadspheres with minimum 6 meshes nd materials per object you quickly end up in a few hundreds or thousands of draw calls required. Ending up in huge CPU load at both the asteroids and planet at close distance, I am pretty sure my performance issues are always the same problem (the draw calls when lots of planes have to be rendered and its CPU impact).

Seems without batching or the MaterialPropertyBlock class supporting computebuffers this can become a dead end in Unity. As most of you pass buffers or lots of values to the shader you probably also create sepate materials per plane.
I wonder if someone has been able to solve this? @NavyFish early mentioned that this can become an issue in Unity.

Has someone somehow workedaround this in Unity?

Good results @JoergZdarsky!

I disappeared with my Normals and uv’s missunderstandings, so i forget about that things - a lot of sheet was optimized and refactored. Noise Engine, and… all of my Renderer.
So i have no problems with normal mapping precision. Idk.
Anyway by some native limitations on unity i’am gonna use own engine. Without bad things as ā€œall core on floatā€.
Anyway it is a good reason for new things to be studied :slightly_smiling:

UPD: Lol. With a lot of beer and ā€œinsomniaā€ i’ve get it!
Quad vertices stitching! Haleloyah! Precumputed and fully procedural. Own inplementation.
Screen

P.s Sorry for my english for last time, please :smile:

1 Like

@JoergZdarsky

The ballpark of recommended draw calls seems to range from ~3k to 15k for a DirectX-9-11 application. Certainly several hundred thousands are likely to affect the frame rate in a non-DirectX-12 application.
I don’t know if Unity allows you to work with buffer types, but according to this document:

procedural planet geometry would run best with UpdateSubresource buffer management.
Map/Unmap would be used for geometry that changes every frame.

I am currently developing a prototype in C++ to first calculate the patches on the CPU, batch them into the buffers and then make the draw calls and doing the batching is kind of tricky, because one has to adapt the index buffer offset for each patch to relate to the patches vertex data.
Hope you find a way to tackle this in Unity.

Ah ok well thank you that gives me some orientation. Could become hard to stay below that once you render a lot on screen, but well thanks to space there is at least lots of empty space :smirk:

Well I reworked the asteroids a bit as they indeed had a few triangles too much :slightly_smiling:
Triangles have been signifcantly pushed down, and the asteroids now use a stone material to give them a bit more ā€œstructureā€. Guess the number of triangles and the look is a good compromise now.

I’ve also experimenting a bit with Worley Noise to give them a bit more natural and uneven look as Worley adds some nice break lines.

Only downside of Worley is that it seems quite computation heavy, more that ~3 octaves take noticable longer to compute. Does anyone know some good alternatives to Worley / Voronoi noise or other very quick different noises besides simplex noise? Otherwise I think I have to decide if these little additional structures are worth the time of additional expensive noise calls. However, it would be nice to add a few craters procedurally at least…

As I’d like to render more asteroids (~500-1000) in a scene I guess I should make the number of vertices per quadtree plane dynamic and depending on the LOD. So that very far away asteroids dont have more that 2 to 4 vertices per plane.

Or what I am also thinking of is eventually even switch to render a single cube instead of separate six planes (the quadsphere) to decrease also the number of draw calls. Or maybe create an Icosphere in the compute shader.
Has one of you already tried to create a simple object (quad, sphere, etc.) in a computeshader?

Well I think the only thing I can do now to decrease triangles where possible and try to simplify farer objects from a quadsphere to another kind of single mesh. Or try to use textures to pass the necessary information to the vertex shader, and don’t use a computebuffer. Best case probably would be Unity would expands MaterialPropertyBlock to support ComputeBuffers. Hopyfully NavyFishs and my request will be noticed somewhen (although there is a low chance I guess).

@JoergZdarsky

Your idea with the LOD-controlled vertex-count per model will help. Pretty much all games make use of it.
Depending on how many different types of asteroids in regards to their shape you have,
you could look into if Unity supports instancing which on DirectX level is drawn via DrawIndexedInstanced.
This could be used for asteroids of the same LOD-level and would make the drawing way more efficient while still allowing size or color variations.
It is described here:
https://msdn.microsoft.com/en-us/library/windows/desktop/bb173349(v=vs.85).aspx
http://http.developer.nvidia.com/GPUGems2/gpugems2_chapter03.html

So. Small progress at the moment.

Edit: So. I decided to work around with DrawProcedural, DrawNow and other procedural drawers…
Result is so incredible!
6 Meshes. 60x60 verts. (64x64 in calculation with border to reduce normal mapping and erosion border stuff)
12 Textures. 2 Per 1 Mesh. 120x120 pixels. (128x128 in calculation with border to reduce normal mapping and erosion border stuff)
Normal map (Normal map + noise map), and color map (Color map + slope map).
Calling Compute shader dispatch every OnRenderObject() for each Mesh.
Compute shader contains 4 cores.
40FPS on GTX 430. Lol :smiley: (I have GTX 580 Rev2 Overlocked on my bookshelf, but my PS need to be upgraded)
And i finally found a magic formula to @NavyFish’s CubeCoord calculation mechanics with LOD.
Up to 12 LOD levels
UNFORTUNATELY! I dont have ā€œnoise deepā€ normal mapping artifacts. Checked on automated noise fractal zooming tech.

Thanks to that thread again!

1 Like

Hey folks, been on a long hiatus and have some catching up to do on the thread. Life got real busy, real quick.

@zameran perhaps I misunderstood, but why are you calling Compute Shader dispatch on every draw? The Compute Shader should be used to build the patch once, and the results stored in VRAM to be reused in subsequent draw calls. Would probably give you a huge boost in performance not calling the Compute Shader every frame. But I may be misinterpreting your usage, so please correct me if I’m wrong.

Will respond to previous posts (and PMs… Thanks/sorry for not responding yet @JoergZdarsky) in due time. Great progress everyone!!

Welcome back! You, guys make me think, that you’re all ignoring me :slightly_smiling:
@NavyFish, I’am new in procedural drawing and operations with GPU. As i said before - my first planet renderer is based only on CPU. All sheet. Only one simple shader with textures for quad. So.

I will work on it(GPU renderer). And i will get new experience on GPGPU calculations, rendering, pipeine, other.

All question that i have - i resolve with google and books. (And with this thread too :wink: )
As a result - misunderstandings. That is normal.

Hello everyone again!
I have some problems with sphere-space (tangent-space) normals.
Texture drawn on top left corner of screen is normal texture for top (Y+) plane (quad).



Camera looks to Z+ planet plane (quad).
So normals calculation provide normal texture per quad.

Code sample
Where …pos.xyz is planet-space position of vertex (already spherized, positioned and height applied).
Can someone push me in good direction?

Anyway, another results with some comments (Bottom left plane is Z+ (front)):
Sobel filter. Calculated from height values of vertexes (already spherized, positioned and height applied)…



As you can see here normals calculated in self-space (yeah?). As a result - wrong directions per quad.

Simple neighbors filter.



The same as a previous…

I cant figure out ā€œthe magicā€.

Thank you in advance.