Procedural Terrain Rendering How-To

Hello Akenre,

First, with respect to your second question regarding applicability to GPU-generated data, yes - you need some terrain height information available to the CPU, as all of these occlusion and LOD tests are performed on the CPU (the host). I have explored performing LOD calculations on the GPU but this is a substantially complex task, and one which I feel the current generation of GL APIs are ill-suited to achieve.

Streaming terrain height data from the GPU to host address space is completely feasible and seamless if done correctly. Unfortunately the low-level GL APIs to achieve the required asynchronous ‘feedback’ are not exposed by the Unity Engine - native plugins are required, none of which currently exist. Joerg keeps hounding me to write one though :slight_smile:

Anyway, you don’t need the full patch height data to determine the appropriate patch-AABB: all you need is the maximum and minimum vertex heights for that patch. This can be approximated through a low-frequency/octave procedural height ‘sampling’ on the CPU (create a low resolution version of the patch host-side), or measured by analysis of the entire patch’s height map. In the latter case, this can be achieved quickly on the GPU by using a parallel scan operation, and then sending back the two results (min and max). Asynchronous transfer is still required in this case, but the memory bandwidth requirement is drastically reduced compared to sending the entire patch’s heightmap and determining min/max on the host.

Back to your original request, I will provide a verbal summary for how I build the transformation matrix - it’s not a simple copy-paste because the calculation takes place in stages throughout multiple parts of the generation system.

Each of my patches has a center point (not actually aligned with a vertex due to using and even-by-even # of vertices per side, i.e. 128x128), and that point is defined by an azimuth and elevation. Each face of the quadsphere defines an azimuthal and elevational rotation axis. For example, the Z+ (normal) face (in a right-handed coordinate system) has an azimuthal axis of +Y, and an elevational axis of +X. All patches belonging to that +Z face reference these axes of rotation.

As an example, the LOD-0 (one patch) for this face has an az/el value of 0 and 0. The LOD-1 patches have the following az/el values:

corner           AZ    EL
--------------------------
top-right        +45   -45
top-left         -45   -45
bottom-left      -45   +45
bottom-right     +45   +45

These rotation values can be applied to a unit vector of the face’s normal vector - i.e. of (0,0,1) for the +Z face. The rotations can be applied in either order, resulting in a unit vector pointing from the planet’s origin to the center of the patch once the patch has been ‘sphericalized’ (all vertices normalized).

This center vector is the first key element of the transformation matrix from world-to-patch space. I usually refer to this as the patch’s ‘up’ vector, or the patch’s ‘normal’ vector (as opposed to an individual vertex’s normal vector).

The patch space’s coordinate system is thus defined:

Patch ‘tangent’ vector: take a unit vector having the same direction as the face’s elevation rotation axis, and rotate it around the face’s azimutal rotation axis by the patch’s azimuth value.

Patch ‘bitangent’ vector: take a unit vector having the same direction as the face’s azimuthal rotation axis, and rotate it around the face’s elevational rotation axis by the patch’s elevation value.

This yields an orthonormalized 3d basis. You can verify this by ensuring Tangent cross Bi-tangent equals Normal.

A standard change of basis matrix may be constructed with this knowledge of the patch’s basis defined using worldspace vectors (technically, these are “planetspace” vectors, which could be further transformed in a multi-planetary scenario).

Of note, when constructing the patch-aligned bounding box, I mathmatically assume the patch has zero curvature. This causes deviation of the BB at the lowest LODs, but I’ve found that it doesn’t cause a problem. At higher LODs, the deviation is imperceptable, since the patch curvature is so minute. Thus, construction of the patch-BB is pretty straightforward, as each of the box’s axis are aligned with the patch basis.

I hope this has helped you understand how I construct the patch AABBs and how to construct a transformation matrix from planet-space to patch-space.

Yes agreed, you most likely need the specific GPU data of that patch on the CPU to perform a propper AABB test.
@Akenre, I’ve tried a inaccurate workaround by considering for the AABB test either absolute possible min/max values (the bounding values of the noise) or some average values, but no matter what you try, you will end up in some inaccurate LOD behavior (planes are culled/splitted/merged too early or too late). So yes, you need some patch information back on the CPU if you do your LOD there as NavyFish said.

This raises an important question as long as you haven’t written the native plugin :wink:.
Does the amount of data significantly matter with regards to the stall happening?
Because I agree with your suggestion that it is absolutely not necessary to read back the full GPU data to the CPU.
Min/Max of the single patch is indeed sufficient for the AABB test.
Or maybe better returning the plane’s bounding coordinates (4x float3) in a separate computebuffer, because then you are able to create a raw collision box with these values.
Hmmm, OK, to see what we talk about I thing I am going to measure that evening what it takes to read back the compete vertice data, only the 4xfloat3 bounding data, or only min/max values.

@Navyfish and @JoergZdarsky
Thanks for the detailed answer, much appreciated. I will test that.
Concerning the read-back from GPU, since this does not have to happen on every frame, maybe
the negative effects might be imperceptible. Even more so if the amount of data is proportional to the effects, as we only need 8 bytes transferred. Even 2*2bytes might work if you try out min16int (-32,768 to +32,767) which suffices for height, but that small difference might be irrelevant.
I work with RWByteAddressBuffer as the output format of the CS for reassignment to the VS at the moment.
Need to test the performance with CPUAccessFlags=D3D11_CPU_ACCESS_READ for the host buffer description instead of 0.
@Navyfish described earlier that an update of the data is only required when the camera has moved
a relevant distance away from the last update’s position.

EDIT: I went ahead and used a 2nd UAV to return the min and max height and suffer a 15% FPS drop when doing it every frame.
The difference originates solely in the data transfer line to map the uav buffer:

dc->Map(g_pBufExtremaResult, 0, D3D11_MAP_READ, 0, &MappedResource);

So if that is done only once in a long interval and cached, it might be viable.
I might stick with with a low-resolution version on the host-side, because i decide the quadtree-segmentation based on the visibility as well, so it seems better to do it host-side before a potential over-segmentation of the quadtree.

Are you guys actually creating 224x224 patches all the way from LOD-0. Or Do you slowly increase the resolution in exponents of 2 up to 224x224 at higher LODs?

Akenre, are you using your own engine? If so, it sounds like you can use asynchronous data access without being hogtied by Unity. Check out this article:

https://msdn.microsoft.com/en-us/library/windows/desktop/bb205132(v=vs.85).aspx#Accessing

Essentially, it recommends you
Create a CPU-accesible buffer with the D3D10_USAGE_STAGING flag, and then execute ID3D10Device::CopyResource to copy the GPU-generated data to the CPU-accesible storage. Then, you wait 3 frames, and finally on the third frame call the appropriate Map method to read the terrain data on the CPU.

This should prevent the performance impact you’re seeing by eliminating the CPU-GPU synchronization, (pipeline stall). If you implement this, let us know your results!

Edit: and yes, i generate patches with the same vertex count at each LOD

@NavyFish
Yes i use straight DirectX11 with C++.
What you linked to is exactly what i am using, except the frames delay.
I will work on something to realize the frame delay. Looks like i will have to
save(copyresource) the buffer within each node and then map it 3 frames later from there.

Where i really see this lag-free-approach potentially shine is to “probe” the whole patch mesh for collision-purposes to map to a corresponding host-representation of the surrounding terrain patches.

Agreed, having the terrain data fully available to the host is really valuable. It enables better LOD and occlusion testing, and allows for complex collision physics with dynamic objects in the world.

Please let us know how adding that frame delay works out!

Oh, wtf, you can’t return the transform feedback?
You can’t output the vertex as texels and retrieve that to the CPU in Unity either? I would think at the very least that’s possible.
I figured since it has an API for procedurally making meshes, you could read something from the GPU to do that.
Welp. Yeah, there’s lots you simply can’t do in Unity. There’s many reasons why I’m not using it and I’m just using OpenGL myself on upcoming projects.

Most of what Unity does is not remotely useful for a Space Game either, especially if when you’re streaming assets. It’s good for scene-based games.
I think you’re generally best off using a general purpose 3d library and other tools for a space game over Unity.

If a developer has access to the Unity source code (or any other framework), he should be able to modify the rendering architecture to realize the read-back. It is however quite a hack to re-use the GPU-address-space buffer from the compute shader in another shader stage without re-transfer even in C++ and it is quite noticeable that this use case has not been sufficiently considered during the design process of current GPUs. At least not out of the box.
A much bigger problem for space games is the CPU overhead most engines make you come up with to tackle the precision issues. In that aspect no engine on the current market can really handle that efficiently (without overhead such as several transformations).
General purpose 3d library and other tools are certainly an option, if you have life-force to implement all the other stuff that engines bring along (Lighting, effects, network code, editors and so forth), should your project require it, that is.
Pretty sure, you will be happier doing it on your own in OpenGL, also a tad more of a challenge hehe.

1 Like

Source code for Unity is quite expensive, as far as I’m aware. You don’t get it with the standard license like UE4. You have to negotiate that with them privately.

@Navyfish:
In post #187
https://forums.inovaestudios.com/t/procedural-terrain-rendering-how-to/765/187
you described that you use normal noise maps for the individual quads of a patch. (I don’t mean the vertex normal value map, but the one that creates the pattern)
Can you elaborate on how you create the map exactly and in which shader?
I assume that would have to be done in stage2 compute shader?

One more question, you said you use

My solution was to create 226x226 vertices with the first kernel.

Did you make the THREADGROUP_SIZE_X/Y then 256 / 32 with 32 ThreadsperGroup for optimal thread usage?

Thank you.

Spent a little time on this project and worked on the depth buffer issue, pretty happy with the results. I made a logarithmic depth buffer implementation in a simple shader. So below are two cubes that are offset by 100 units and have a scale of 10k units, the camera has a far clipping plane of 1mil units, the gif shows what happens with a standard Unity shader vs the custom one.

Will have to do some research and see how to make this work with a standard unity shader, a pre-pass would be nice, anyone have experience with this or ideas?

2 Likes

Nice @cybercritic! Though I cant help here or do any suggestions, I faced that zbuffer issue too of course, but didnt work on a solution myself yet.

Meanwhile I stumbled across yesterdays newsletter in Unity, which goes the Vulkan way in its actual beta 5.5.4. So I am getting curious if 5.5 will include some new stuff usable.

1 Like

New Screenshots, im working on Texturing…
https://dl.dropboxusercontent.com/u/19373509/slope.PNG
https://dl.dropboxusercontent.com/u/19373509/slope2.PNG

Anybody have an idea how i generate nice mountains and Islands / Continent? What is the best way to combine Noise Functions? im doing it still so:
`

>     float f = ridgedmf(position, 16, 1.5, 1.02, 1.2);
		float test = rmf(position, 18, 13.885, 1.85, 0.6)*2;
		float res = f * test;
		float ve = voronoi(position, 16) * 3;
		float xx = turbulence(position, 231231, 16)*ve;
		res += 1000;
		res -= 998;
		res = max(-3, res+ xx);
		res = getTerraced(res, 0.14f, 3);
		xx = max(0, res);
	   return float4(res, 0, 0, 1);`

and im not shure if my normals are correct… When i look to the left the color are more red:
https://dl.dropboxusercontent.com/u/19373509/normalLeft.PNG

When i look to the right, the color are more green:

https://dl.dropboxusercontent.com/u/19373509/normalRight.PNG

I use this Function to generate the normals on the GPU (From a heightmap):

float2 uv = input.UV;
uv.y = 1 - uv.y;

float _TexelSize = 1.0 / PatchWidth;

float2 tn = uv + float2(0, -_TexelSize);
float2 ts = uv + float2(0, _TexelSize);
float2 te = uv + float2(_TexelSize, 0);
float2 tw = uv + float2(-_TexelSize, 0);

float2 cc = uv;

float n = tex2D(heightmapSampler, tn.y < 0 ? cc : tn).r;
float s = tex2D(heightmapSampler, ts.y > 1 ? cc : ts).r;
float e = tex2D(heightmapSampler, te.x > 1 ? cc : te).r;
float w = tex2D(heightmapSampler, tw.x < 0 ? cc : tw).r;

float ew = e - w;
float ns = n - s;

float3 result = float3(ew *NormalScale, ns * NormalScale, 2 * _TexelSize);

result = normalize(result);
  return float4(result* 0.5 + 0.5, 1.0);

in the PlanetSHader i do:

	float3 normal = tex2D(NormalMapSampler, input.UV).rgb;
	normal = mul(normal, World);

hm…

5 Likes

@montify

Looks good. Are you using slope for texture determination?

About your normal calculation, it looks alright, just not so sure about this line?

return float4(result* 0.5 + 0.5, 1.0);

the * 0.5 + 0.5 is done why?

Question:
From your previous videos and the shader syntax, i assume you work with DirectX and your own engine.
Are you using precomputed atmospheric multiple scattering or the simpler single scattering variant?
I am having a real hard time translating the OpenGL implementation to DirectX for multiple scattering.

Hello, *0.5 +0.5 is to translate in a 0 -1 range?! I dont really know… :confused:

Yes i use XNA/DirectX …
I use the approach by Sean O’Neil: Atmospheric Scattering… the other is from Proland Precomputed Scattering the Oneil one looks fine for me.

Which approach do you mean?

For Slope: i have a 250x1 texture with different color, than i use the normal like:
> float slope = 1.0f - normal.b;

 diffuse = tex1D(gradientSampler, 0.45 + slope);

Here a Video… i have some cracks between the 6 sides, somtimes… and the Atmospheric scattering shader still not complete… now i miss the sun :frowning:
https://www.youtube.com/watch?v=IzZ0c5bWRIU&feature=youtu.be

3 Likes

@Montify

To check the normals, you could try to put a simple moveable pointlight or simple directional light on an untextured planet patch.
Then when moving it around hills , you can see the normals better. With the scattering shader it’s difficult to see in your video.
Another option is what @NavyFish has shown in post #187 , a visualization of the normal vectors.
[Procedural Terrain Rendering How-To]

This can be done with relatively little effort using an additional drawcall with a geometry shader stage outputting lines from point primitive input comprising of the patch vertex data.

Have you tried omitting the line *0.5 +0.5 ? Just an idea.

Hello, yes but than the normals are brocken, becasue out of 0-1 range i guess…
i assume that the normals are correct for now :wink:

Anybody experience with Terrain Shadows?

Anyone also using the Unity5.5 for experimental purposes?
They included a few new functions I feel they could be intersting for us. E.g. exposed meshes and computebuffers in their 5.5 roadmap (currently in beta). Unfortunately their isn’t too much information around yet.

Beta: 5.5
Download Beta
Stabilization in progress. Target : Nov 2016

Exposed Meshes and ComputeBuffers to native code plugins [Recently updated]
Graphics
Mesh.GetNativeIndexBufferPtr
Mesh.GetNativeVertexBufferPtr
ComputeBuffer.GetNativeBufferPtr

I am improving the vertex normal part of my engine currently. Actually also went ahead and built the normal visualization.
My approach of calculating the vertex normals is also in the 2nd CS-stage, but not from a height map, instead directly from the stage 1 noise z-values of the vertices using the difference algorithm. The problem is that this simple algorithm produces the normals in patch space, but not oriented correctly in regards to the sphere. So to convert those to world space, i would have to rotate those normal vectors by the angles given by the spherical normalized vertex position of stage 1 to orient them correctly. To do so there are a couple of quaternion transformations needed, which are kind of heavy. I wonder if i am overcomplicating this somehow though and a more lightweight approach is existing. Would appreciate some input on how others go about and do that.

The normal map that i saved out from the compute shader for debugging purposes looks strange, as if it was 90 degrees sidewards of the patch. This is the 2DTexture of the normal map:

I cannot explain that, the method used is :

int inBuffOffset = id.x + id.y * nVerticesPerSide;

float n = asfloat(patchInput.Load(((inBuffOffset - nVerticesPerSide) * byteStride) + heightIndex));
float s = asfloat(patchInput.Load(((inBuffOffset + nVerticesPerSide) * byteStride) + heightIndex));
float e = asfloat(patchInput.Load(((inBuffOffset + 1) * byteStride) + heightIndex));
float w = asfloat(patchInput.Load(((inBuffOffset - 1) * byteStride) + heightIndex));

float ew = e - w;
float ns = s - n;

float3 normalVector = normalize(float3(ew, ns, 2 * spacing));

normalMap[id.xy] = float3(normalVector.x , normalVector.y , normalVector.z );

This is really hard to debug, maybe someone has an idea where the mistake is?

EDIT: Doh! The ominous operation posted by montify improved it:

normalVector = normalVector* 0.5 + 0.5;
normalMap[id.xy] = float3(normalVector.x , normalVector.y , normalVector.z );

Unfortunately the Visual Studio Graphics Analyzer is buggy for me and does not show any values read out of a raw buffer in a compute shader, so i was not aware about the numeric range of the normals, which is addressed by the multiplication with 0.5 and adding 0.5.

There is still something not right i think with the large area more red colorations in the lower half. Nearly as if the patch curvature is included and adds to the red somehow.
With all 3 color channels, it looks like this:

EDIT2: Th reason for the wrong coloration seems to have been that i took the actual terrain height for the real-world-sized patch for the difference algorithm. When using the final noise value only and assuming a spacing of 2*1, the normal map now looks better:

1 Like