I went through this a few months ago in Rust. I wrote all the code by hand, no LLMs. Then I went ahead and added a small "game" on top, plus some special effects like pixelization shaders and chromatic aberration at the edge of a flashlight.
if anyone is interested! The repo has lots and lots of in-progress screenshots so you can see the renderer come to life, plus all the hilarious visual bugs along the way.
I learned a lot! My biggest lesson, other than the specifics of how rendering works, was that modern CPUs are really fast: a single-threaded CPU renderer can definitely run an interactive 3D game with some fancy special effects.
Yes, exactly this -- it's a fast and convenient way for my code to just write the pixels into a spot in RAM (the CPU's RAM, not VRAM) and have those pixels end up on screen. On modern architectures this nearly always goes through a GPU, so even if you're not using the GPU to accelerate your 3D rendering/math, you gotta deal with it just to put pixels on a screen at the end of the day. So that's what wgpu does for me.
This resource, along with Mathematics for Computer Graphics by John Vince [1], was truly indispensable when I wrote my own software renderer [2]. This was long before LLMs, so the whole process took me at least a couple months - most of it trying to wrap my head around math behind computer graphics and tracking down C segmentation faults. Fun times.
I wish we could have just one of these tutorials properly cover the concern of triangle clipping. This is the part that I struggle with the most in a software renderer. If you are going to be building a practical one, this is something you will eventually have to deal with, even for super basic scenes. Any time geometry intersects the view frustum you need to clip those triangles.
Sorta unrelated but I really enjoy (I don't want to use the past tense as I often refer back to it) your article on Emulator Backed Remakes (https://gabrielgambetta.com/remakes.html). Thanks!
Thanks for saying that :) Have you built something based on that idea? I still haven't, but I'm thinking on throwing some Claude at it for the most mechanical parts.
You only need to clip triangles is you're worried about attribute interpolation for very large triangles. There's two ways to handle this: (1) discard (fast but not a great user experience); or, (2) primitive synthesis. Just frustum clipping is enabled by point picking in the local tile. Primitive synthesis requires some FP kung fu; but, is easiest done in barycentric space against a reverse transformed clipping rectangle. This lets you carefully control clip rounding error using either doubles or (better) fixed point. Abrash likes to use integer fixed point, but that is historical — modern fixed point can be handled with careful control of the fp unit in the mantissa. The major issue is regenerating the Z and the 1/Z values for the new vertices of the synthesized primitives. Everything else should flow down the pipe naturally, assuming a deferred attribute synthesis rasterizer.
There are examples in the open source version of my rasterizer: OpenSWR.org.
I'm clearly super old-school when it comes to rasterisation, and a lot of this has flown above my head. I have a ton of questions; I hope you can answer them.
> (1) discard (fast but not a great user experience)
What are we discarding here, and why is it fast but not a great user experience?
> (2) primitive synthesis
I assume this is retriangulating clipped triangles that are now no longer triangles?
> reverse transformed clipping rectangle
Which spaces does this reverse transformation map from and to? I assume the clipping rectangle here is the triangle's AABB in raster space (or as you say, barycentric space).
> integer fixed point, but that is historical — modern fixed point can be handled with careful control of the fp unit in the mantissa
So we are no longer doing 16.16 fixed point, but tweaking the FP representation itself?
> The major issue is regenerating the Z and the 1/Z values for the new vertices
Why is this a major issue?
> deferred attribute synthesis rasterizer
I assume this means attributes are perspective-correct interpolated in raster space.
Sorry I was so terse! Let's ignore guard band, for now. Our rasterization surface is a rectangle. We don't need to "geometrically clip" a triangle to the rectangle's surface. Instead, we just walk the surface of the rectangle and ask two questions: (1) does the triangle capture this (sub)sample; and, (2) what's the interpolated value of the attributes at this (sub)sample. In practice, for software rasterizers, we're working on tiny subrectangles (the tile), e.g., 4x4 or 2x8, whatever. So, we can be a little "inefficient" with our walking with respect to the edge testing for interiority. (This is roughly how HW works at the 2x2 level, as well.) Because the (sub)sample is "pulling" the attribute, we don't need to geometrically "clip" the triangle to the tile: the tile's (x,y) subsamples do that "for free".
On the flip side, if a triangle is 'really big' we need a guard band to either reject or subdivide triangles. The first one is fairly cheap -- we're throwing away the triangle! -- the second one is a better user experience, but requires synthesizing primitives. (The worst case is that a single triangle becomes 5 triangles, I think.) Each of those triangles needs its Z and 1/Z calculated in fixed precision. The precision of that fixed precision (though) can be clamped to the local tile; so, even though the global precision might need to be 25.25 (or whatever), the tile-local precision is only 4.9 (or whatever), with an intermediate 24.24 that can be handled with a float-float patch-up. The computation should all occur in the triangle's barycentric space: that means you need the inverted barycentric mapping to invert the guard band into the triangle's barycentric space. You do that because it lets you control the fixed point calculations better. (You can leave off the inverted determinant multiplication until the last moment.)
When I say "deferred attribute synthesis" I mean that we don't calculate attributes in the vertex shader. Instead, we calculate the barycentric, Z, and 1/Z values and pass those along. When we fire up the tile walker for the triangle (in general we only need 1-3 tiles), we calculate the attributes "on the fly, as they're used" and then let the compiler do CSE to fold down the replicated constructions.
My own knowledge of GPU rasterization may be dated, but IIRC GPUs tend to rely on guard band clipping up to a guard band threshold before using geometric clipping. The guard band clipping involves rejecting 2D coarse rasterization blocks that are fully outside of the scissor rect. This is just a quick rectangle check, but the tradeoff is that a larger guard band means more GPU time lost in over-rasterization and potential higher precision requirements for rasterization values (which could be fixed point). Beyond the guard band, the triangles are clipped in floating point against the frustum clip planes.
My renderer attempts always got stuck on the "should implement clipping" phase too, until I finally bit the bullet and managed to write a working one without much effort, independently "rediscovering" the Sutherland–Hodgman algorithm [1] as I found out later (googling it beforehand would've been cheating, of course).
The algorithm itself is fairly straightforward and intuitive, I think the biggest mental block is the weirdness of the projective space and working with homogeneous coordinates (actually the only frustum plane that you have to clip against in P₃(ℝ) is the front plane, the rest could be clipped after the perspective division, but no reason not to do it all at the same time while you're at it). The plane equations in the clip space are super simple, basically the six equations of the form ax + by + cz = w simplify to
x = ±w
y = ±w
z = ±w.
Meaning, for example, that if the x coordinate of your vertex is greater than the w coordinate, that vertex is outside the right clipping plane. The Sutherland–Hodgman itself goes something like this:
# Returns true if point is inside the half-space defined by plane
def point_inside_plane(point, plane) -> bool:
# single dot product, can be further simplified
# Returns t such that the edge (p1, p2) intersects plane at lerp(t, p1, p2)
def edge_intersect_plane(edge: (Point, Point), plane) -> float:
# single dot product, can be further simplified
# Given the vertices of a simple polygon and a plane,
# returns the part of the polygon fully inside the plane
def clip_against_plane(poly: [Vertex], plane):
let result: [Vertex] = []
let [(v_1, v_2), (v_2, v_3), ..., (v_n, v_1)] = poly.edges()
for each (v_i, v_j) of the edges:
let i_inside = point_inside_plane(v_i, plane)
let j_inside = point_inside_plane(n_j, plane)
if i_inside and j_inside:
# v_j will be pushed on the next iteration!
result.push(v_i)
else if not i_inside and not j_inside:
pass # Nothing to do!
else:
# One is inside, the other is not, we have to clip
let t = edge_intersect_plane((v_i.pos, v_j.pos), plane)
# Synthetize a new vertex straddling the plane
let v_new = Vertex(
pos = lerp(t, v_i.pos, v_j.pos),
# For each vertex attribute
attrib = lerp(t, v_i.attrib, v_j.attrib)
)
if i_inside:
result.push(v_i); result.push(v_new) # discard v_j
else:
result.push(v_new); result.push(v_j) # discard v_i
return result
Then you just call this for all the planes so that the output of one call becomes the input for the next call! The end result of this process is a convex polygon (of at most nine vertices for a triangle against six planes), which can be trivially triangulated. You can make the whole process faster by precomputing so-called outcodes which allow you to avoid clipping triangles known to be entirely outside at last one plane, or entirely inside every plane.
This is a great comment. As someone who is familiar with R^3 geometry but not projection or clipping, I feel like I could almost implement it from your comment alone.
Can you explain more about clipping before/after the divide by w operation? Is this about avoiding coordinates with w = 0? Otherwise all the geometric ideas seem to make sense in both the world/view-space frustum and the ±1 box.
There was a software rendered game in vein of Tomb Raider 1 graphics for both DOS and Unix, but I can't remember its name. It was an exploration game, modern, a bit cyberpunkish, in 3D.
POVray is a declarative language used to render raytraced scenes.
This is a tutorial that demonstrates how the rasterization technique works from a programmatic perspective.
https://github.com/kshitijl/tinyrenderer-rs
if anyone is interested! The repo has lots and lots of in-progress screenshots so you can see the renderer come to life, plus all the hilarious visual bugs along the way.
I learned a lot! My biggest lesson, other than the specifics of how rendering works, was that modern CPUs are really fast: a single-threaded CPU renderer can definitely run an interactive 3D game with some fancy special effects.
(I originally wrote this a few years ago but other people maintain it now)
[1]: https://www.amazon.co.uk/Mathematics-Computer-Graphics-John-...
[2]: https://github.com/nkanaev/tipsy
There are examples in the open source version of my rasterizer: OpenSWR.org.
> (1) discard (fast but not a great user experience)
What are we discarding here, and why is it fast but not a great user experience?
> (2) primitive synthesis
I assume this is retriangulating clipped triangles that are now no longer triangles?
> reverse transformed clipping rectangle
Which spaces does this reverse transformation map from and to? I assume the clipping rectangle here is the triangle's AABB in raster space (or as you say, barycentric space).
> integer fixed point, but that is historical — modern fixed point can be handled with careful control of the fp unit in the mantissa
So we are no longer doing 16.16 fixed point, but tweaking the FP representation itself?
> The major issue is regenerating the Z and the 1/Z values for the new vertices
Why is this a major issue?
> deferred attribute synthesis rasterizer
I assume this means attributes are perspective-correct interpolated in raster space.
On the flip side, if a triangle is 'really big' we need a guard band to either reject or subdivide triangles. The first one is fairly cheap -- we're throwing away the triangle! -- the second one is a better user experience, but requires synthesizing primitives. (The worst case is that a single triangle becomes 5 triangles, I think.) Each of those triangles needs its Z and 1/Z calculated in fixed precision. The precision of that fixed precision (though) can be clamped to the local tile; so, even though the global precision might need to be 25.25 (or whatever), the tile-local precision is only 4.9 (or whatever), with an intermediate 24.24 that can be handled with a float-float patch-up. The computation should all occur in the triangle's barycentric space: that means you need the inverted barycentric mapping to invert the guard band into the triangle's barycentric space. You do that because it lets you control the fixed point calculations better. (You can leave off the inverted determinant multiplication until the last moment.)
When I say "deferred attribute synthesis" I mean that we don't calculate attributes in the vertex shader. Instead, we calculate the barycentric, Z, and 1/Z values and pass those along. When we fire up the tile walker for the triangle (in general we only need 1-3 tiles), we calculate the attributes "on the fly, as they're used" and then let the compiler do CSE to fold down the replicated constructions.
The algorithm itself is fairly straightforward and intuitive, I think the biggest mental block is the weirdness of the projective space and working with homogeneous coordinates (actually the only frustum plane that you have to clip against in P₃(ℝ) is the front plane, the rest could be clipped after the perspective division, but no reason not to do it all at the same time while you're at it). The plane equations in the clip space are super simple, basically the six equations of the form ax + by + cz = w simplify to
Meaning, for example, that if the x coordinate of your vertex is greater than the w coordinate, that vertex is outside the right clipping plane. The Sutherland–Hodgman itself goes something like this: Then you just call this for all the planes so that the output of one call becomes the input for the next call! The end result of this process is a convex polygon (of at most nine vertices for a triangle against six planes), which can be trivially triangulated. You can make the whole process faster by precomputing so-called outcodes which allow you to avoid clipping triangles known to be entirely outside at last one plane, or entirely inside every plane.[1]: I. Sutherland and G. Hodgman. 1974. "Reentrant polygon clipping." Communications of the ACM, Volume 17, Issue. Available: https://dl.acm.org/doi/10.1145/360767.360802
Can you explain more about clipping before/after the divide by w operation? Is this about avoiding coordinates with w = 0? Otherwise all the geometric ideas seem to make sense in both the world/view-space frustum and the ±1 box.
Personally I’d have gone with a teapot as my demo but I don’t see anything shocking about the model the author chose.
mbappe'?
An error occurred: API rate limit already exceeded "