As an update to the render issues I had with versions after 5.2.4, it was related to the RenderTexture-Format. Unity seems to have changed the default bevahiour when working with those, so take care what you do when using RenderTextures that you explicitely define the right format.
I’ve been further looking into the latest 5.4 beta (5.4b18) and the MaterialPropertyBlock. Unity included SetBuffer() to the MaterialPropertyBlock in that beta branch. So the good news is, you can now pass your ComputeBuffer results via MaterialPropertyBlock to your renderer, which I and a few others (NavyFish) did to pass ComputeShader results (terrain heigth values) via ComputeBuffer to the Vertex Shader (to replace the vertices there). And I can confirm that this works as intended.
However the bad news is, it will still break batching. Even when using a single mesh instance and a single material instance. Each time you change the MaterialPropertyBlock by doing a new SetTexture() or SetBuffer() call, batching will be prevented. So this alone won’t help to increase performance for those who have chosen a similar design to create terrain in Unity on the GPU.
Unity also added “GPU instancing” into the 5.4 beta branch, but looking into the current documentation and dicussion in the beta forums, different MaterialPropertyBlocks will prevent instancing on GPU too at the current stage of implementation. It will be of help to render lots of similar objects which werent batched before, which can be of use in our scenario to render suns (or maybe small repetitive asteroids/rocks). I did a quick brute force test where I created ~25.000 spherical meshes (without LOD or anything) which weren’t batched before. The GPU Instancing feature worked out of the box here.
So the good news is it looks like Unity’s 5.4 branch will invent a new feature for those who want to create not only planets but more some kind of universe where you draw repetitive but lots of objects.
However, I am afraid when drawing a detailed terrain surface where you need different appearances per plane by use of the ComputeShader, there won’t be any new options in terms of performance gain.
I have some new experience with Unity 5.4b18 last week, but my planetary engine crashes the driver on quad generation/dispatching, so i rolled back. I think that it is Unity’s bug, cuz no crashes on build.
Anyway i tested instancing too - 10k cubes without lod etc. ~ the same results to you.
BTW, @JoergZdarsky, in your demonstration videos i can’t see any “render stall” when new quads generated. What is your dispatching mechanism?
At the moment i use WaitForEndOfFrame (several) to “smooth” the “render stall” on dispatching.
Ho do you do the magic?
Thank you.
No need for WaitForEndOfFrame or anything, I simply dispatch everything to the GPU for calculation and directly render it afterwards in my shader. Do you read back anything to the CPU?
No. All is on GPU!.. oh, whait…
So first of all i pass my generation parameters to compute shader (Create buffers, pass generation constants, quad generation constants)
After that i call ComputeShader.Dispatch(…) for all threads that i need.
In the end, in rendering queue, i use Graphics.DrawMesh(…), and pass in to it default prototype mesh (simple flat quad mesh) with output ComputeBuffer that contains position data for each vertex.
In quad vertex shader i simply move vertices to it positions (taken from passed output ComputeBuffer)
I can post more detailed and commented code here…
In another words:
Pack parameters/buffers to ComputeShader. (On quad creation)
Call ComputeShader.Dispatch(…) after packing. (Only once per quad life cycle!)
Pack output buffer from dispatched ComputeShader in to the quad shader. (In Update)
Graphics.DrawMesh(prototypeMesh, …) (In Update)
If generation parameters was changed - restart from p.1.
Without WaitForEndOfFrame(several times eg. 8 or 4) i have Gfx.WaitForPresent up to 100-300ms.
I think i lost something and i need to draw mesh directly from the GPU? Like Graphics.DrawProcedural?
Thats basically the way most of us do it (at least NavyFish and me also). No idea where these stalls come from, except that you may have the same problem with the number of draw calls per frame. Or some other kind of overhead or flaw in your code (e.g. too many computeshader dispatches per frame, bad computeshader thread sizing, or something in the computeshader itself…). Or maybe hardware. I dont know.
Here I dont have any stalls (GTX 570), except that the number of draw calls remain a problem, as stated before.
A little bit further brute force testing of Unity3Ds GPU instancing. 10.000 asteroids within a little more dense volumetric fog, all rotating around the origin and the fixed camera. No care for triangle count or anything.
Looks like this is going to be usefull, I think I am going to use that feature in my next iteration for the asteroid fields and procedural asteroids on GPU. As soon as 5.4 is out I might rework my engine a bit and start on reimplementing parts of the universe creation (new placement of suns and planets, asteroid fields, and reworks of parts of floating origin and scaled spaces)
Unity 5.4 finally has been released a few days ago. I am curious to see if others are trying out some of the new additions that now went into their final as well?
Two things I found interesting are
Graphics: GPU Instancing Support
Use GPU instancing to draw a large amount of identical geometries with very few draw calls.
Works with MeshRenderers that use the same material and the same mesh.
Only needs a few changes to your shader to enable it for instancing. Supports custom vertex/fragment shader and surface shaders.
Set per-instance shader properties from script via MaterialPropertyBlock.
Supports Graphics.DrawMesh command.
Graphics: Improved multithreaded rendering:
Compared to current dual-thread rendering (main thread + rendering thread), this splits up rendering logic into concurrent “graphics jobs” that run on all available CPU cores.
My idea is to initially experiment with these changes with a new asteroid implementation. As far as from the beta’s limitations this requires a new strategy as far as I can see, as GPU Instancing didnt work work with different MaterialBlockProperties applied - so the strategy of using ComputeBuffers and utilizing the GPU still doesnt work
If you create a larger group of different asteroids procedurally at the initial startup and then performance friendly in the background during gameplay, and render asteroids out of this group instead of calculating them completely individual on demand, GPU instancing could help a lot nonetheless. I think this is the way Limit Theory did this too.
Thats why I currently try to make my head around what would be the best mesh strategy for asteroid using GPU instancing. Icospheres, being a single mesh and gameobject, would probably the best and GPU friendliest. You instance only once per LOD, and it forms for lower lods a nice unique sphere. However a large downside IMHO is the quickly increasing number of vertices during the splitting iterations. Maybe this is not a problem when instancing, but IMHO gives you less flexibility on refining the LOD (and limits the number of vertice detail at the highest LOD to something Unity3D can handle in a single mesh).
The other idea would be to keep going with quadspheres, which would be much more flexible. But as long you want to use GPU Instancing you’d have to instance at least six meshes for a single asteroid until you combines the meshes to one mesh.
Are others here looking into the new Unity 5.4 version as well to share experiences?
I have not revisited this project in some time, but that planet rendering itch is starting to bother me again, so I’m very thankful you’ve been keeping up with the latest unity advancements and reporting on them. Will save me and others a lot of trial and error!
I’m also on the verge of a major life transition: after 12.5 years wearing a uniform, I will be taking it off once and for all at the end of is month. I’m very excited for this transition. I may need to find a new screenname, however! lol
Point is, I will soon have a lot of free time, and, while there are many ways to fill it that I currently have planned, this is a project I will likely revisit sooner than later.
I intend to work on the gpu -> cpu readback stalls. Likely this will require me to build a native plugin to utilize asynchronous buffer transfers. I think it can be done.
Have you considered contacting the unity development regarding SetBuffer breaking batches? I was under the impression that batching was the whole point of material property blocks, so it’s possible that that’s a bug.
In closing, again thanks for the great updates. I look forward to doing a little more active work on this project in the near future.
glad you are back and that you are doing well!
Sounds like a big turning point in your live, wish you all the best for that and that you enjoy that new part of your life a lot!
Well thanks, however I wouldn’t be overoptimistic with the new Unity3D version 5.4. There seem to be some very interesting features in it pretty usable for various usecases for universe stuff, but for some of the planetary surface issues we faced here I have doubts that there is something revolutionary new in. But I still have to look into some of the details to figure out how to use this or that best. I’ve not been putting too much time into this yet, as continouing on a beta version can be a bit anoying. Now as 5.4 went into the final version, I’ve recently upgraded to 5.4 with my latest project, which is now a third iteration of the prototype I am building from scratch (well, still utilizing lots of the techniques and classes previous built, but I felt it is time to do that to structure the creation of space objects more efficiently). And besides that writing a few blog articles/tutorials so that the next one does not have to go through some of the same pains. I look really forward to rebuilt some of the parts, e.g. the handling of scales and management of multiple objects when it comes to LOD, and make some things more natural (better denser asteroid fields, or placing planets more realistally around sun). We’ll see. And, as the new 5.4 features are stable, it now makes sense to use them too.
That would be a huge step forward!! For us here on the forum for our specific procedural planet usecase of course, but I have seen many requests on the Unity3D forum as well. Being able to get GPU data back to CPU would allow to implement many features we are currently blocked. I you find your time and manage to do this, a lot of people would owe you something thats for sure.
Well… I fear that we mght have overrated the MaterialPropertyBlocks a bit. My current impression is, looking what blocks the batching, that MPBs are more meant for a more elegant way to pass the same material properties (textures, buffers, floats, etc) to different objects at once instead to set each one seperately. I hope I am wrong, but that is how it looks to me.
I’ve been in contact with one of the developers (zeroyao) describing our very specific problem that different buffers break batching. This has been the response.
Firstly I’m sorry that the previously mentioned DrawMeshInstanced API
has been postponed for 5.5. There was some interdependencies to the
other team’s work and after we came back from HackWeek it was too late
for 5.4.
And for your particular problem I think the best way to do is to set the
big ComputeBuffer and the texture to your shared material, and have
some instanced properties in each MaterialPropertyBlock for reference
different location for the buffer and texture (e.g. a UV offset for the
texture).
Currently textures and computebuffers can’t be instanced automatically.
Only float values (scalar, vector or matrix) There is this new
TextureArray API you can try out, but some manually texture packing code
is required. ComputeBuffers can never be an array. Use a big combined
buffer and use instanced properties for location as mentioned above.
Hmmm. Sounds reasonable and could work. As you see from his comment, not only the buffers are a problem to batching, different materials such as textures are as well. Anyway maybe there is an opportunity in his suggestion. I have a raw layout of a dynamic management/pooling of multiple shared textures and buffers in my head which I plan to refine and implement in the next couple of weeks.
Besides that implementation I am going to see if the Texture2DArray is a helpful alternative to the atlas and if it helps for batching. Aras had a good talk on that topic here: http://forum.unity3d.com/threads/texture-arrays.392028/
Same here, looking forward to some interesting and fun discussions!
EDIT: I finished implementing a “shared RenderTexture” management/pooling. Right now on request of a new RenderTexture, a slot of a configurable larger texture (basically an atlas) is being provided including indivual UV coordinates. If all slots of a shared texture are used a new one is being created in the background. Including house-keeping (use the next free slot and release shared textures again if not used).
The dimensions (or number of slots) of the shared textures in the background are configurable in width and height. E.g. in case you use 256x256 textures for normalmaps, those are being stored for example on a 1024x8192 texture in the background, which means it offers 128 slots.
I have to combine this now with our current GPU/ComputeShadee-based procudural planet/asteroid implemenation. If this basically works (the texturing works with these shared textures and individual UV coordinates), then I’ll expand the shared concept to ComputeBuffers too. At that time then, the magical batching should start to work in Unity fingers-crossed
I combined the shared texture strategy to the procedural generation (our CS/GPU approach) in Unity3D. I am testing this with asteroid rendering currently.
According to the debug log the request of new shared texture slots and the background management of these shared textures works, also the pixel- and UVOffset values per slot look good. Also the surface shader now only uses a part of the texture (e.g. 64x64, the height and width of each slots texture) instead of using the full shared texture (e.g. 64x1024). So far so good.
Last remaining change is to consider the offset values in the ComputeShader and SurfaceShader. In the picture you see that I am currently write (in the CS) and read (in the SS) only into the first slot as the texture repeats per all planes using the same shared texture instance (I currently only constantly use offset (0,0) for testing purposes). Thus that is the expected result at the moment. As soon as I include the offset values passed to the ComputeShader and SurfaceShader things should look good. Then 16384 planes could share the same texture in 64x64 resolution or 4096 planes in 128x128 resolution.
That part should be finished in the next couple of days, then I can move over to do the same for ComputeBuffer.
If things work well, the current asteroid should be rendered in one batch instead of thousands.
Wow, great progress! This is a cool technique indeed, reminds me almost of John Carmack’s ‘supertexture’ technique. Assuming it works for you, I definitely will incorporate the same approach into my project once I get going again. Thanks for the update!
There is some progress with the shared RenderTextures. I added the offset values to the ComputeShader and SurfaceShader as planed. It basically worked immediately but nevertheless, I had a bad time the last few days because there was one annoying bug. At the moment a quadtree node was splitted, sudden black texturing appread on other planes. Guessing there was a bug with the offset calulcation in the SharedTexturesManager class, I was investigation there first.
In the end the issue was quite simple. I had one RenderTexture.Release() command in my quadtree code left that was called at the point a quadtree node is splitted. Of course this may not be done when working with a shared RenderTextures that functions as an atlas. The one and only class that may finally release the Texture is the SharedTextureManager that validates if a texture is used anymore or not.
Works like a charm now, I am pretty happy with the SharedTexture implementation. Quite flexible by defining the number of slots in both directions width and height, and the background housekeeping removes textures if not used in order so safe memory.
Only thing I now need to get under control before moving on to do the same with ComputeBuffers is the bleeding that now happens due to the atlas. I haven’t yet dealed with the problem in detail, need to check if padding within the shared RenderTexture is the solution or if I should handle this problem somewhere else (e.g. modifing the UVOffset calculation in the SharedTexture class or later in the SurfaceShader slightly). Any suggestion very welcome!
While implementing the slot strategy for ComputeBuffers I am currently investigating on how to create slots for the buffer and read from them or the right offset in the vertex shader.
I am currently stuck at how the vertex shader has to be changed.
In C# the count is increaded by the number of slots: this.computeBuffer = new ComputeBuffer(count * slots, stride , computeBufferType);
The ComputeShader writes into the offset (with the offset constant being the slot ID as an int, so 1,2,3,…n) int outBuffOffset = (id.x + id.y * constants.nVertsPerEdge) * constants.sharedPatchGeneratedFinalDataBufferOffset; patchGeneratedFinalDataBuffer[outBuffOffset].position = float4(patchCoordCentered,1);
Question is how to change the vertex shader to read here from the right offset of the computebuffer too… float4 position = patchGeneratedFinalDataBuffer[v.id].position;
I wonder, especially as the shared RenderTexture approach now works nicely, if I might move onto storing the results of the ComputeShader in a texture instead of a ComputeBuffer to pass the height information etc. from there to the Vertex-/SurfaceShader? While ComputeBuffers are still fine to pass parameters to the ComputeShader, storing its result (height-information and position information) might be not necessary or a texture would do the same work?
Are there reasons against using RWTexture to store heightmap etc. in a texture instead of a buffer and do a lookup in the vertex shader? I’d then create two textures, one that holds the vertex position (float4) and one that holds the patchCenter (float3). I’ve read that ATI cards had once problems with that. I keep looking to do the slot mechanism for the computebuffer, but would like to as about the textures meanwhile.
I did my own itty bitty procedural gen 5 years ago.
It had issues, very incomplete, and all on the CPU. And Unity didn’t do doubles back then which was screwing me. Only spent about 2 weeks on it, IIRC, as a learning exercise.
All I did is learn how to procedurally cubespheres plot and manipulate a cubespheres, equally space the UV and polygons(so corners aren’t pinched), rewrite a Java Simplex noise implementation to C#, and experiment with generating and blending that noise. Super basic.
It sounds like you may need to pass the slot ID to the vertex shader as a constant input in order to pick from the compute buffer. Material Property Blocks should let you do so, but perhaps I misunderstood the situation.
Nothing wrong with writing to a RW texture per-se, especially since your data is inherently 4x4-byte aligned. Compute buffers are a bit more flexible when it comes to the data types they store but in this case you don’t need that flexibility. I would still profile both methods to make sure there isn’t some huge performance difference. Nice work with the shared text you! I will have to consider implementing this method myself.
Awesome! Thank you for sharing. I love how full this community is of individuals who are compelled to build a procedural planet renderer, and who tend to describe it as a ‘learning exercise’ and ‘super basic’
Oh no, this thread is being bumped again, already made me look at my code a few days ago. Anyway, don’t think I posted this video, it’s from the last PG spell.
Passing the slot ID to the vertex shader is not the issue. I currently set a uniform int “_uniform int PatchGeneratedFinalDataBufferOffset;” from outside with the slot ID. Moving everything to the MaterialProperty would be my last step then as soon as I get the ComputeBuffer running.
My problem is to read from the right index of the buffer in the vertex shader.
While the change in the compute shader to write there seemed obvious… int outBuffOffset = (id.x + id.y * constants.nVertsPerEdge) * constants.sharedPatchGeneratedFinalDataBufferOffset;
… I wonder how to change the following two lines in the vertex shader that read from the buffer, currently using
uniform int _PatchGeneratedFinalDataBufferOffset;
void vert(inout appdata_full_compute v, out Input o)
{
UNITY_INITIALIZE_OUTPUT(Input, o);
#ifdef SHADER_API_D3D11
// Read Data from buffer
float4 position = patchGeneratedFinalDataBuffer[v.id].position;
float3 patchCenter = patchGeneratedFinalDataBuffer[v.id].patchCenter;
#endif
}
v.id or SV_VertexID is AFAIK created under the hood, but somehow that must be changed to consider the slot ID value stored in _PatchGeneratedFinalDataBufferOffset too. Simply multiplying that value was too easy thinking, the asteroid looks pretty scrambled afterwards
Yes that my goal currently. The implementation of the slots is done, seems the vertex shader indexing is the only issue left, thus I dont want to put that aside now yet.
@innociv Nice one! Impressed that you were able to turn away from the project once you’ve gotten that far already
I am sorry ;). The dynamic cloud movement is nice, definitely another topic we haven’t scratched here yet. Did you apply some certain rules to define the wind directions of each segment?
It’s very simple rules for the wind direction atm, I just sample a few points around the wind marker and add a direction offset based on the height at those sampled points. Suppose this could be extended to any data that is available at the sample point like temperature, pollution or precipitation.
About your problem, off the top of my head, just round the texture position to an int in your shader and use that for the computed buffer index, you might need an offset, depending on the orientation difference between texture and buffer though.
well not followed the discussion but just throwing an idea regarding double precision planet rendering, i have also restarted my work on this tech after getting bored with it in jan this year. I made my planet to the point where I struggled with double precision and collision issues as I was throwing a lot of work to gpu and kept cpu mostly free. But I have always struggled with this approach with planets as transfering data from gpu to cpu is a bottleneck and is a requirement for a lot of tasks, so my solution that I have restarted working on and I am sure some one may have discussed it in this thread above is to load everything related to terrain geometry generation on to cpu but divide the task on a separate thread, so one thread does the task of rendering whereas the other thread calculates/streams the data for the new / reused nodes (memory pooling/caching). This way you can scale the planets with double precision noise generation of terrain and/or also add custom height maps with circular blending on the planet surface on cpu and only generate the normal map on gpu. The 2nd thread also removes the hiccups in frame rate. It will also help with collision detection. You can also replicate this with planets oceans if you want to simulate floating physics. I have successfully made multi threaded approach work with 22 levels of detail without any frame rate hiccups on cpu, but my algo is still work in progress. so just throwing an idea for some one struggling with double precision planets.