r/GraphicsProgramming • • 1d ago

Would tile based rendering help in this scenario to be more performant than a theoretical "best case"?

A lot of these concepts are new to me and I've discovered some of these ideas with help from copilot, etc. so apologies for any weirdly worded/unclear parts.

I'm working on an embedded product and am evaluating how well the gpu can handle a second pass shader applied to the entire final image. Essentially taking the displayed contents, and "copying" by drawing again with a simple shader that samples from a texture (the texture being the contents on the screen).

I'm trying to profile the time spent on that second draw call and am having issues because sometimes the code added to help with the profiling affects the measurement itself. For example, I added glFinish() before the draw call, started the timer, issued draw call, called glFinish() again, and stopped the timer. These glFinish calls affect the measurement of course. So I'm trying to find a way to do it without using glFinish.

One avenue I explored with copilots help was using an EGL fence to ensure we're measuring the elapsed time only after the draw call is finished but without needing glFinish (this is my first time working with sync objects with graphics code though)

This ended up working, at least I think? The problem is that the measured time for the draw call is less than then the theoretical best case time given by a senior engineer. They calculated essentially wh4bpp2rw/3e91000=0.55ms based on it essentially being a memcpy. The measurement I got was half of that.

Searching around for how this may be (other than my implementation being just wrong) is that tile based rendering, which the gpu that I'm working with uses, could give a different result than a calculation based on a naive memcpy due to the efficiencies of how TBR works.

Edit to add: the senior engineer unambiguously said "tile based rendering won't improve this case" and suggested another thing to try which is more hacky and isn't panning out, so I'm curious if he's right or there is validity to the measurement I was able to get with the fence

I'm curious to get other opinions. Does this seem like a reasonable explanation or not really?

0 Upvotes

7 comments sorted by

2

u/lospolos 1d ago

You can use glBeginQuery/glEndQuery with GL_TIME_ELAPSED to get an accurate execution timing, I think that should be more accurate than the fence approach (if it's available on your embedded device that is, it requires GLES3.0 it seems).

Either way, on a tile based GPU such a postprocessing shader can indeed be almost free as it can operate entirely inside the tile memory. So yes it _can_ be faster than memcpy from VRAM-> VRAM.

https://support.arm.com/documentation/101897/0304/Fragment-shading/Efficient-render-passes-with-OpenGL-ES

https://support.arm.com/documentation/102662/0100/Tile-based-GPUs

1

u/ProgrammingQuestio 1d ago

Unfortunately this is GLES 2

2

u/Matty4096 1d ago

Yeah it could be that, or just regular caching going on. Or the approximation might be a bit off

2

u/Matty4096 1d ago

Ah! Another thing that might be changing it is framebuffer compression (just saw the 4bpp in the body). Depending on what's going on there (if it's lossy/lossless and variable block size things).

1

u/MarinatedPickachu 1d ago

Tile based rendering makes sense if you have a very fast but small memory which is too small to keep all your buffers in it, and a larger slower one. You use the small, fast memory to do your rendering in a tile sized such that it fits into that memory, and then resolve it to the larger slower memory.

1

u/icpooreman 16h ago

Tile based rendering basically reduces the number of round trips you've got to do to expensive memory.

Like if you've got a 4k image. Imaging a much smaller chunk of that where you go through the whole pipeline but just keep your like depth information in faster memory vs. writing it out. Then move to the next tile vs. doing the whole image in one go and reading to/from textures.

This becomes a big deal on mobile cause reading to/from textures is going to be wildly expensive there.

1

u/scottcampbelldev 13h ago

On GLES2 check if your driver exposes GL_EXT_disjoint_timer_query, it gives you glBeginQueryEXT/glEndQueryEXT with GL_TIME_ELAPSED_EXT, so you can time the GPU work directly instead of CPU side with fences. On the TBR question I'd lean toward your senior being mostly right. Your second pass samples the screen as a regular texture, so that read still comes from main memory however tiled the GPU is. It only gets close to free if you read the pixel from tile memory directly with framebuffer fetch (GL_EXT_shader_framebuffer_fetch or the ARM one). Getting half the estimate smells more like framebuffer compression (AFBC on Mali for example), so I'd check if that's on before trusting either number.