State Estimation and Bayesian Filtering

This chapter discusses the problem of statistical filtering, namely, the class of problems in which data from one or more noisy sensors are available. These data represent observations of the dynamic state of a system that cannot be observed directly but must be estimated. The procedure used to obtain the best estimate of a system's internal state is called “filtering”, since it removes the various noise components. The evolution of a system (that is, the evolution of its internal state) must follow known physical laws, which are affected by a noise component (process noise). Knowledge of the equations governing the evolution of the state makes it possible to obtain a better estimate of the internal state.

A physical process can be represented in state space (State Space Model) by a function describing how the state $\mathbf{x}_t$ evolves over time:

\begin{displaymath}
\dot{\mathbf{x}}_t = f(t, \mathbf{x}_t, \mathbf{u}_t, \mathbf{w}_t)
\end{displaymath} (3.1)

where $\mathbf{u}_t$ denotes any known system inputs, and $\mathbf{w}_t$ denotes the parameter representing the process noise, that is, the randomness governing the evolution. Likewise, the observation of the state is a process affected by noise, in this case called observation noise. Here too, it is possible to define a function modeling the observation $\mathbf{z}_t$ as
\begin{displaymath}
\mathbf{z}_t = h (t, \mathbf{x}_t, \mathbf{v}_t )
\end{displaymath} (3.2)

where $\mathbf{v}_t$ is observation noise and a function only of the current state.

This formalism is described in the continuous-time domain. In practical applications, signals are sampled at discrete times $k$, and a discrete-time version is therefore normally used in the form

\begin{displaymath}
\begin{array}{l}
x_{k+1} = f_k(x_k, u_k, w_k) \\
z_k = h_k(x_k, v_k) \\
\end{array}\end{displaymath} (3.3)

where $w_k$ and $v_k$ can be regarded as white-noise sequences with known statistics.

Figure 3.1: Example of the evolution and observation of a Markovian system.
Image fig_markov

In systems satisfying equations (3.3), state evolution depends only on the previous state, whereas observation depends only on the current state (Figure 3.1). If a system satisfies these assumptions, the process is said to be Markovian: the evolution and observation of the system depend only on the current state and not on past states. Access to information about the state always occurs indirectly through observation (Hidden Markov Model).

Many approaches for estimating the unknown state of a system from a set of measurements do not account for the noisy nature of these observations. It is possible to construct an algorithm that performs nonlinear regression on the observations to estimate all the states in the problem by solving an optimization problem with a large number of unknowns.

Unlike regression methods, filters aim to provide the best estimate of the state variables as observation data become available. From a theoretical standpoint, regression provides the optimal result, whereas filtering converges to the correct result only after a sufficiently large number of samples.

Bayesian filters aim to estimate, at the discrete time instant $k$, the state of the random variable $\mathbf{x}_k \in \mathbb{R}^n$ given an indirect observation of the system, $\mathbf{z}_k \in \mathbb{R}^{m}$.

Filtering techniques provide both the best estimate of the unknown state $\mathbf{x}_k $ and the multivariate probability distribution $p(\mathbf{x}_k)$ representing the available knowledge of the state itself.

Given an observation of the system, it is possible to define a probability density for $\mathbf{x}_k $ a posteriori of observing the event $\mathbf{z}_{k}$, based precisely on the additional information provided by that observation:

\begin{displaymath}
p^{+}(\mathbf{x}_k) = p(\mathbf{x}_k \vert \mathbf{z}_k)
\end{displaymath} (3.4)

where, as a conditional probability, $p(\mathbf{x}_k \vert \mathbf{z}_k)$ denotes the probability that the hidden state is $\mathbf{x}_k $ given the observation $\mathbf{z}_k$. The “function” $p(\mathbf{x}_k \vert \mathbf{z}_k)$ represents the state measurement model (measurement model). In the literature, the posterior distribution $p^{+}(\mathbf{x}_k)$ is also referred to as the belief.

Applying Bayes' theorem to equation (3.4) yields

\begin{displaymath}
p(\mathbf{x}_k \vert \mathbf{z}_k) = c_k p(\mathbf{z}_k \vert \mathbf{x}_k) p(\mathbf{x}_{k})
\end{displaymath} (3.5)

where $c_k$ is the normalization factor such that $\int p(\mathbf{x}_k \vert \mathbf{z}_k)=1$. Knowledge of $p(\mathbf{z}_k \vert \mathbf{x}_k)$ is essential; it represents the probability that the observation is precisely the quantity $\mathbf{z}_k$ observed given the possible state $\mathbf{x}_k $. The use of Bayes' theorem to estimate the state given the observation is why this class of filters is called Bayesian.

In addition to the a posteriori knowledge of the probability distribution, further information can be used to improve the estimate: knowledge a priori of the observation, obtained from the constraint that the state does not evolve in a completely unpredictable manner but can instead evolve only in certain ways with certain probabilities. These possible modes of system evolution depend only on the current state. The Markov process assumption implies that the only past state affecting the evolution of the system is the state at time $k-1$, namely, $p(x_k \vert x_{1:k-1}) = p(x_k \vert x_{k-1})$.

It is therefore possible to perform an a priori prediction using the Chapman–Kolmogorov equation:

\begin{displaymath}
p^{-}(\mathbf{x}_k) = \int p(\mathbf{x}_k \vert \mathbf{x}_{k-1}, \mathbf{u}_{k}) p(\mathbf{x}_{k-1}) d\mathbf{x}_{k-1}
\end{displaymath} (3.6)

where $p(\mathbf{x}_k \vert \mathbf{x}_{k-1}, \mathbf{u}_{k})$ represents the system dynamics (dynamic model) and $\mathbf{u}_{k}$ denotes any inputs that influence the evolution of the system and are known completely.

Given the a priori state estimate and the observation $z_k$, equation (3.4) can be rewritten as the state update equation

\begin{displaymath}
p^{+}(\mathbf{x}_k) = c_k p(\mathbf{z}_k \vert \mathbf{x}_k) p^{-}(\mathbf{x}_k)
\end{displaymath} (3.7)

The state is estimated by alternating a prediction phase (a priori estimation) with an observation phase (a posteriori estimation). This iterative process is called recursive Bayesian estimation (Recursive Bayesian Estimation).

The techniques described in this section use only the latest available observation to estimate the state. Formally, the discussion can be extended to the case in which all observations are used to obtain a more accurate state estimate. In this case, the filtering and prediction equations become

\begin{displaymath}
\begin{array}{l}
p(\mathbf{x}_k \vert \mathbf{z}_{1:k} ) = ...
...mathbf{x}_k \vert \mathbf{z}_{1:k} ) d \mathbf{x}_k
\end{array}\end{displaymath} (3.8)

For simplicity and because of the lower computational cost, normally only the latest observation is evaluated. In some cases, however (for example, in particle filters), knowledge of the entire past history can be incorporated into the equations relatively easily.

Since continuous variables are being estimated, Bayesian theory cannot be used “directly”. Several approaches have therefore been proposed in the literature to enable efficient estimation, both computationally and in terms of memory usage.

Depending on whether the problem is linear or nonlinear and whether the noise probability distribution is Gaussian or non-Gaussian, each of these filters performs more or less optimally.

The Kalman filter (Section 3.2) is optimal when the problem is linear and the noise distribution is Gaussian. The Extended Kalman and Sigma-Point Kalman filters, Sections 3.4 and 3.5, respectively, are suboptimal filters for nonlinear problems with Gaussian noise distributions (or distributions close to Gaussian). Finally, particle filters are a suboptimal solution for nonlinear problems with non-Gaussian noise distributions.

Grid-based filters (Section 3.1) and particle filters (Section 3.8) operate on a discrete representation of the state, whereas Kalman, Extended Kalman, and Sigma-Point filters operate on a continuous representation of the state.

Kalman, Extended Kalman, and Sigma-Point Kalman filters estimate the uncertainty distribution (of the state, process, or observation) as a single Gaussian. Multimodal extensions such as Multi-hypothesis tracking (MHT) make it possible to apply Kalman filters to distributions such as Gaussian mixtures, whereas particle and grid-based filters are inherently multimodal.

An excellent survey of Bayesian filtering is (Che03).



Subsections
Paolo medici
2026-10-06