Statistical Mechanics of Semantic Compression

Can, Tankut

Condensed Matter > Disordered Systems and Neural Networks

arXiv:2503.00612 (cond-mat)

[Submitted on 1 Mar 2025]

Title:Statistical Mechanics of Semantic Compression

Authors:Tankut Can

View PDF HTML (experimental)

Abstract:The basic problem of semantic compression is to minimize the length of a message while preserving its meaning. This differs from classical notions of compression in that the distortion is not measured directly at the level of bits, but rather in an abstract semantic space. In order to make this precise, we take inspiration from cognitive neuroscience and machine learning and model semantic space as a continuous Euclidean vector space. In such a space, stimuli like speech, images, or even ideas, are mapped to high-dimensional real vectors, and the location of these embeddings determines their meaning relative to other embeddings. This suggests that a natural metric for semantic similarity is just the Euclidean distance, which is what we use in this work. We map the optimization problem of determining the minimal-length, meaning-preserving message to a spin glass Hamiltonian and solve the resulting statistical mechanics problem using replica theory. We map out the replica symmetric phase diagram, identifying distinct phases of semantic compression: a first-order transition occurs between lossy and lossless compression, whereas a continuous crossover is seen from extractive to abstractive compression. We conclude by showing numerical simulations of compressions obtained by simulated annealing and greedy algorithms, and argue that while the problem of finding a meaning-preserving compression is computationally hard in the worst case, there exist efficient algorithms which achieve near optimal performance in the typical case.

Comments:	20 pages (including appendix), 3 figures
Subjects:	Disordered Systems and Neural Networks (cond-mat.dis-nn); Statistical Mechanics (cond-mat.stat-mech); Computation and Language (cs.CL); Neurons and Cognition (q-bio.NC)
Cite as:	arXiv:2503.00612 [cond-mat.dis-nn]
	(or arXiv:2503.00612v1 [cond-mat.dis-nn] for this version)
	https://doi.org/10.48550/arXiv.2503.00612

Submission history

From: Tankut Can [view email]
[v1] Sat, 1 Mar 2025 20:38:16 UTC (99 KB)

Condensed Matter > Disordered Systems and Neural Networks

Title:Statistical Mechanics of Semantic Compression

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Condensed Matter > Disordered Systems and Neural Networks

Title:Statistical Mechanics of Semantic Compression

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators