> This scan is made possible by recent advances in Gaussian Splatting. This is an emerging technology that lets us quickly create very detailed models just from photographs. For this model (or splat, as we call them), my friend Daylen and I flew our drones around Sutro Tower at a respectful distance for an afternoon until we had collected a few thousand photographs.
> I then aligned these pictures in free software called RealityCapture. Alignment is the process that teaches the computer that a bunch of points in different images all actually correspond with the same point in real life. Then I used another piece of free software called gsplat to produce the 3D model itself.
I'm assuming this was done similarly. Very cool.
It's like seeing the promise of Google Street View realized.
The technology is in such an early stage and it already slaps this hard.
Wonder what the next steps are for this tech? "Un-splatting" optimizations to convert things like flat surfaces into classic rigid textured geometry? Factoring the lighting away from the geometry? Filling the pipeline with AI tooling - segmentation to remove the "variables" like humans automatically, generative AI to fill in the missing non-semantic details based on the surroundings, more generative AI to "port" the data from things like public high resolution images of plaques or other undercaptured details into the scene?
Seeing this kind of reconstruction now makes me very excited for what the 3D environments of the future are going to be like.
A lot more architecture would benefit from being available like this splat model. The Pharonic tombs which you can visit in google/other virtual tourism are good, but this is also good!
I noticed that complex patterns - marble, some paintings, some stained glass - are rendered in what seems like stretched, hazy ellipses. Is that the fundamental building block of a splat?
Also, I was interested if any specularity would be captured by the splat. Or any other aspect of the building which shifts from viewpoint - like lights reflected on the marble floor, or where a metal element outdoors is reflecting the sun but only from one vantage point. I didn’t see that. As I understand it, splats are a kind of solution to the “problem” posed by the series of photos so they are capable of capturing that sort of thing. Is that present here, and I missed it? Or is that something that you need a lot of processing power for, to generate a more complex solution for the light fields?
The model seems to be having a hard time with the marble floor in the quire (low resolution), is this because of its reflection?
Firefox, Windows 11, Nvidia
Get the following in the console when it stops working:
Uncaptured WebGPU error: Buffer binding 2 range 138024504 exceeds `max_*_buffer_binding_size` limit 134217728
Uncaptured WebGPU error: Buffer binding 1 range 138024504 exceeds `max_*_buffer_binding_size` limit 134217728
Uncaptured WebGPU error: In a set_bind_group command, caused by: BindGroup with '' label is invalidAnd whenever I see videos explaining the Gaussian splats, they always claim they are supposed to fix lighting issues.
Is this a function of "not enough detail" or.. Something else?
The cutaway Peek feature is like a sci fi version of an Stephen Biesty's Incredible Cross-Sections book. Incredible work!
- let us say i wanted to bring a city to life like this
- what does the procedure look like on my end
- what inputs, tools, libraries would i need to go about doing this?
Unrelated silly question. Are you the same Vincent Woo who went over the whole unit combination story on Medium of a supervisor?
I bet there's a story there.
It's unfortunate that this tour makes it very difficult to see. Really not a fan of this "on rails" presentation style. I understand the angles matter when viewing a splat, but come on.