A short 9-page introduction to the main concepts and results of information geometry for beginners.
The Jensen-Shannon divergence (JSD) is a symmetrization of the Kullback-Leibler divergence bounded by log 2 (even when supports differs) but is not available in closed-form for Gaussians. Instead of taking the arithmetic mean mixture M(p,q)=(p+q)/2, we consider generic means M and define the M-JSD(p,q)=KL(p:M(p,q))+KL(q:M(p,q)). In particular, the geometric G-JSD between Gaussians or the harmonic H-JSD between Cauchy distributions are available in closed-form. Follow-up work
Consider the discrete distributions p=(p1,...,p(n+1)) on the standard simplex (categorical distributions including Bernoulli for n=1) as a mixture family (statistical model of order n). Jensen-Shannon divergence between p and p' rewrites as a Jensen divergence between w=(p1,...,pn) and w'=(p1',...,pn'), and we solve JSD centroid by using the difference of convex algorithm (DCA/CCCP). Allows to cluster robustly histograms with potentially empty bins.
Give arbitrary lower and upper bounds for calculating the Fisher-Rao distance between multivariate normal distributions.
Voronoi diagrams with respect to Bregman divergences which are asymmetric (except for squared Euclidean distances) are studied: Bregman VDs can be computed by clipping power diagrams or by using the lifting to the Bregman potential function (generalizing the paraboloid lifting transform).
By building a 1D exponential family for any two probability measures dominated by a given measure, we show that the log-normalizer is the negative of the Bhattacharrya divergence (hence concave). Prove that the best Chernoff parameter is the intersection of an exponential geodesic with a dual mixture bisector.