Learning maximum entropy models from finite-size data sets: A fast data-driven algorithm allows sampling from the posterior distribution

Ulisse Ferrari

doi:10.1103/PhysRevE.94.023301

Article Dans Une Revue Physical Review E Année : 2016

Learning maximum entropy models from finite-size data sets: A fast data-driven algorithm allows sampling from the posterior distribution

(1)

Ulisse Ferrari

Fonction : Auteur
PersonId : 750045
IdHAL : uferrari
ORCID : 0000-0002-3131-5537
IdRef : 279962215

Institut de la Vision

Résumé

Maximum entropy models provide the least constrained probability distributions that reproduce statistical properties of experimental datasets. In this work we characterize the learning dynamics that maximizes the log-likelihood in the case of large but finite datasets. We first show how the steepest descent dynamics is not optimal as it is slowed down by the inhomogeneous curvature of the model parameters' space. We then provide a way for rectifying this space which relies only on dataset properties and does not require large computational efforts. We conclude by solving the long-time limit of the parameters' dynamics including the randomness generated by the systematic use of Gibbs sampling. In this stochastic framework, rather than converging to a fixed point, the dynamics reaches a stationary distribution, which for the rectified dynamics reproduces the posterior distribution of the parameters. We sum up all these insights in a “rectified” data-driven algorithm that is fast and by sampling from the parameters' posterior avoids both under- and overfitting along all the directions of the parameters' space. Through the learning of pairwise Ising models from the recording of a large population of retina neurons, we show how our algorithm outperforms the steepest descent method.

Domaines

Physique [physics]

Fichier principal

1507.04254.pdf (1.46 Mo)

Origine	Fichiers produits par l'(les) auteur(s)

Gestionnaire HAL-UPMC : Connectez-vous pour contacter le contributeur

https://hal.sorbonne-universite.fr/hal-01379105

Soumis le : jeudi 15 décembre 2022-16:56:07

Dernière modification le : lundi 18 mars 2024-15:36:03

Dates et versions

hal-01379105 , version 1 (15-12-2022)

Identifiants

HAL Id : hal-01379105 , version 1
ARXIV : 1507.04254
DOI : 10.1103/PhysRevE.94.023301
PUBMED : 27627406

Citer

Ulisse Ferrari. Learning maximum entropy models from finite-size data sets: A fast data-driven algorithm allows sampling from the posterior distribution. Physical Review E , 2016, 94 (2), pp.023301. ⟨10.1103/PhysRevE.94.023301⟩. ⟨hal-01379105⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSERM UPMC CNRS U968 SORBONNE-UNIVERSITE SU-MEDECINE

60 Consultations

54 Téléchargements

Learning maximum entropy models from finite-size data sets: A fast data-driven algorithm allows sampling from the posterior distribution

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager