3D Gaussian Splatting

The idea behind 3D Gaussian Splatting is to represent the scene using a set of three-dimensional Gaussian primitives. 3D Gaussians are based on the three-dimensional extension of one-dimensional Gaussians. Three-dimensional Gaussians are defined by a covariance matrix $\Sigma$ (in world coordinates) and centered at the point (mean) $\mu$:

\begin{displaymath}
G(\mathbf{x}) = e^{-\frac{1}{2} \left( \mathbf{x} - \mu \right)^{\top} \Sigma^{-1} \left( \mathbf{x} - \mu \right) }
\end{displaymath} (10.116)

(the normalization factor is omitted because the function is used as the primitive's profile during rendering).

To be rendered, this Gaussian must first be transformed into camera coordinates through a rotation $\mathbf{W}$ and then projected into image coordinates. However, one can use an approximation in which a two-dimensional Gaussian is drawn in image space. By locally linearizing the perspective projection through its Jacobian $\mathbf{J}$, the three-dimensional covariance is projected into a two-dimensional covariance $\Sigma'$:

\begin{displaymath}
\Sigma' = \mathbf{J} \mathbf{W} \Sigma \mathbf{W}^{\top} \mathbf{J}^{\top}
\end{displaymath} (10.117)

where $\mathbf{W}$ is the rotational part of the transformation alone, and where, as an approximation, the Jacobian $\mathbf{J}$ of the perspective projection is evaluated at the point $(x,y,z)^{\top}$ transformed by the camera rotation and translation. For example, in the case of a pinhole camera projection:
\begin{displaymath}
\mathbf{J} = \begin{bmatrix}
k_u / z & 0 & - \frac{k_u x}{ z^2 } \\
0 & k_v / z & - \frac{k_v y}{ z^2 } \\
\end{bmatrix}\end{displaymath} (10.118)

The matrix $\Sigma'$ is therefore a $2 \times 2$ (ZPvBG01) matrix that describes the covariance of the projected Gaussian ellipse in the image plane.

In (KKLD23), a further step is taken: since parameterizing a covariance matrix (positive semidefinite) is difficult, one starts from the fact that the matrix $\Sigma$ represents an ellipsoid and can therefore be minimally parameterized instead of treating all the matrix entries as unknowns. The idea is to use a scale matrix $\mathbf{S}$ (3 DOF) and a rotation matrix $\mathbf{R}$ (another 3 DOF, but normally represented by a quaternion; see Section A.3):

\begin{displaymath}
\Sigma = \mathbf{R} \mathbf{S} \mathbf{S}^{\top} \mathbf{R}^{\top}
\end{displaymath} (10.119)

thus parameterizing the covariance of each Gaussian with 6 DOF 10.2. Note that $\mathbf{S} \mathbf{S}^{\top} = \diag \left( s_x^2, s_y^2, s_z^2\right)$.

Finally, each 'Gaussian' may be associated with an RGB color or spherical harmonics (Spherical Harmonics, SH), in addition to the opacity parameter $\alpha$, similar to that used in NeRF. In practice, the Gaussians are rendered from nearest to farthest until the opacity saturates.



Footnotes

... DOF10.2
The complete primitive is normally described by a three-dimensional position, opacity, covariance parameters, and color coefficients.
Paolo medici
2026-10-01