3D Gaussian Splatting

The idea behind 3D Gaussian Splatting is to represent the scene using a set of three-dimensional Gaussian primitives. 3D Gaussians are based on the three-dimensional extension of one-dimensional Gaussians. Three-dimensional Gaussians are defined by a covariance matrix $\Sigma$ (in world coordinates) and centered at the point (mean) $\mu$:

\begin{displaymath}
G(\mathbf{x}) = e^{-\frac{1}{2} \left( \mathbf{x} - \mu \right)^{\top} \Sigma^{-1} \left( \mathbf{x} - \mu \right) }
\end{displaymath} (10.119)

(the normalization factor is omitted because the function is used as the primitive's profile during rendering).

Before it can be rendered, this Gaussian must first be transformed into the camera coordinate system and then projected onto the image plane. Since perspective projection is nonlinear, a local approximation is introduced by linearizing the projection using its Jacobian $\mathbf{J}$. The three-dimensional covariance is thus projected into a two-dimensional covariance $\Sigma'$:

\begin{displaymath}
\Sigma' = \mathbf{J} \mathbf{W} \Sigma \mathbf{W}^{\top} \mathbf{J}^{\top}
\end{displaymath} (10.120)

where $\mathbf{W}$ represents only the rotational component of the camera-to-world transformation. Translation does not directly affect the covariance, but it determines the scene point $(x,y,z)$ in camera coordinates at which the Jacobian $\mathbf{J}$ of the perspective projection is evaluated. For example, in the case of a pinhole camera projection:
\begin{displaymath}
\mathbf{J} = \begin{bmatrix}
k_u / z & 0 & - \frac{k_u x}{ z^2 } \\
0 & k_v / z & - \frac{k_v y}{ z^2 } \\
\end{bmatrix}\end{displaymath} (10.121)

The matrix $\Sigma'$ is therefore a $2 \times 2$ (ZPvBG01) matrix describing the covariance of the Gaussian ellipse projected onto the image plane.

In (KKLD23), a further step is taken: because parameterizing a covariance matrix (positive semidefinite) is difficult, one starts from the fact that the matrix $\Sigma$ defines a family of level ellipsoids. It is therefore possible to use a minimal parameterization based on the dimensions and orientation of this ellipsoid, rather than treating all the matrix elements separately. The idea is to use a scale matrix $\mathbf{S}$ (3 DOF) and a rotation matrix $\mathbf{R}$ (another 3 DOF, but normally represented by a quaternion; see Section A.3):

\begin{displaymath}
\Sigma = \mathbf{R} \mathbf{S} \mathbf{S}^{\top} \mathbf{R}^{\top}
\end{displaymath} (10.122)

thus parameterizing the covariance of each Gaussian with 6 DOF 10.7. Note that $\mathbf{S} \mathbf{S}^{\top} = \diag \left( s_x^2, s_y^2, s_z^2\right)$.

Finally, each Gaussian may be associated with an RGB color or with a direction-dependent radiance representation using spherical harmonics (Spherical Harmonics, SH), in addition to the opacity parameter $\alpha$, analogous to that used in NeRF. In practice, the Gaussians are rendered in depth order, from nearest to farthest, progressively compositing their radiometric contributions10.8.



Footnotes

... DOF10.7
The complete primitive is normally described by a three-dimensional position, opacity, covariance parameters, and color coefficients.
... contributions10.8
The process can be terminated early when the residual transmittance becomes sufficiently small.
Paolo medici
2026-10-06