Thank you very much for your reply @INovaeFlavien / Flavien .
Iâve been checking google and my short research resulted in the same (only limited math in combination with doubles) like you mentioned. And even if possible I do not really want to restrict my implementation only to most current GPUs. However eventually I will try to add some simple double values and do some basic math into the ComputeShader to see at least if the compiler complains or not and eventually try to debug some results.
May I ask if you utilize the GPU to calculate your planetary patches, or at least for parts of the process? Not asking for details, just a simple yes or no (probably yes) so that I can have at least a little hope that it might make sense to investigate into that further more on my own.
tl/dr: Initial thoughts about the float precision issue and maybe how to overcome it.
So I set a breakpooint at my patch creation for a level 18 node to check the data of the node.
node.Level = 18
node.OuterVector1 = (-0.1113586,1,-0.4103546)
node.OuterVector2 = (-0.111351, 1,-0.4103546)
node.OuterVector3 = (-0.1113586,1,-0.4103622)
node.OuterVector4 = (-0.111351, 1,-0.4103622)
node.CenterVector = (-0.1113548,1,-0.4103584)
node.Diameter = 1.078959E-05
patchConstants.spacingPixels = 2.99192E-08
patchConstants.spacingVerts = 1.211015E-07
patchConstants.nPixelsPerEdgeWithSkirt = 258
patchConstants.nVertsPerEdge = 64
This means that the math to create the vertice and pixel positions initally in [-1|+1] cube space
on GPU (before normalizing and pushing the data out of [-1|+1] to real world radius) has to handle the following situation (which is currently happening on GPU in float3):
OuterVector1(x) -0,1113586
OuterVector2(x) -0,111351
Distance (x) -7,6E-06
Split by 256 -2,96875E-08
Indeed this looks like being dangerously close to float precision issues already here at the very early point within my node data.
As my approach of using a cubesphere typically created in a [-1|+1] cube first, then normalized and applying radius, and organizing this in a quadtree, I made a table to check how precise I am currently with that.
Looks surprisingly not toooo bad for a test implementation just looking at the resolution, but not optimal, especially when you think about rendering bigger planets.
So where am I, asuming we wonât be able to change the full GPU workflow in the ComputeShader to double precision?
The process currents consists of:
0.0 Check/Update Quadtree Node / Create patch Parameters | Resource-Impact: LOW
1.0 UV-Coordinate and Triangle-Incides arrays | Resource-Impact: LOW
2.1 Initial Vertex/Pixel Position arrays [-1|+1] - flat cubespace | Resource-Impact: LOW
2.2 Updated Vertex/Pixel Position arrays [-1|+1] - normalized | Resource-Impact: LOW
2.3 Updated Vertex/Pixel Position arrays [realworld] - multiplied by radius | Resource-Impact: LOW
3.0 Noise array (SimplexNoise using Vertex/Pixel Position arrays[realworld]) | Resource-Impact: HIGH
4.0 Updated Vertex/Pixel Position arrays [realworld] - add/subsctract by noise | Resource-Impact: LOW
5.0 Normals and NormalMap-Texture (by calculating normals out of Vertex/Pixel Position arrays [realworld] input) | Resource-Impact: MEDIUM-HIGH
6.0 Material-Texture (using noise data and normals as input) | Resource-Impact: MEDIUM
with:
0.0 Check/Update Quadtree Node / Create patch Parameters (float) (CPU)
CPU -> GPU
1.0 UV-Coordinate and Triangle-Incides arrays (float) (CPU)
2.1 Initial Vertex/Pixel Position arrays [-1|+1] - flat cubespace (float) (GPU)
2.2 Updated Vertex/Pixel Position arrays [-1|+1] - normalized (float) (GPU)
2.3 Updated Vertex/Pixel Position arrays [realworld] - multiplied by radius (float) (GPU)
3.0 Noise array (Precision: float) (GPU)
4.0 Updated Vertex/Pixel Position arrays [realworld] - add/subsctract by noise (float) (GPU)
5.0 Normals and NormalMap-Texture (float) (GPU)
6.0 Material-Texture (float) (GPU)
GPU -> CPU (asynch)
As the precision issue seems to be originating in 0.0 within the node data being floats, and asuming the GPU can only handle float, my initial idea would be to change the process by moving certain parts back to the CPU and change it to double. Basically the CPU would calculate the round spherical surface and send it to the GPU to calculate noise (step 2.3->3.0) and normals.
At this point where send positional data from CPU to GPU I would have to downgrade double back to float again for the ComputeShader (but probably keep my double precision data on the GPU).
The GPU would use this downgraded float as input data to calculate noise and furthermore the normals and normalmaptexture.
0.0 Check/Update Quadtree Node / Create patch Parameters (double) (CPU)
1.0 UV-Coordinate and Triangle-Incides arrays (double) (CPU)
2.1 Initial Vertex/Pixel Position arrays [-1|+1] - flat cubespace (double) (CPU)
2.2 Updated Vertex/Pixel Position arrays [-1|+1] - normalized (double) (CPU)
2.3 Updated Vertex/Pixel Position arrays [realworld] - multiplied by radius (double) (CPU)
CPU -> GPU
3.0 Noise array (float) (GPU)
4.0 Updated Vertex/Pixel Position arrays [realworld] - add/subsctract by noise (float) (GPU)
5.0 Normals and NormalMap-Texture (float) (GPU)
6.0 Material-Texture (float) (GPU)
GPU -> CPU (asynch)
The interesting question is, will there be a good way to cast the double back to float on the CPU before dispatching it to the GPU so that it can act as valid input for the noise algorithm there. And furthermore, would the downgraded spherical position data be sufficient, after applying the noise result to them on the GPU to continue to work with that data (downgraded position data and noise result)
to calculate the normals, without reading the noise back to the CPU inbetween.
Maybe one solution would be to multiply the spherical [-1|+1] positions (in double) not by the radius but by some value to shift numbers and that prevents the loss of digits when casting to float, and send these temporary positions to the GPUâs noise function to work with that instead of the realworld ones (that become might become too big for a float when getting closer to the surface).
Another solution might be to create two temporary position arrays sent to the GPU, one HIGH one that holds the double data before the digits (so xxxxxxx.0) and one LOW one that holds the data after the digits (so 0.xxxxxx), call the GPUâs noise function twice with both variants, and stack the result.
Curious about anyones thoughs, has someone of you thought about preventing loss of precision in a planetary scenario already?