Matching of Markov Databases Under Random Column Repetitions

Bakirtas, Serhat; Erkip, Elza

Computer Science > Information Theory

arXiv:2202.01730 (cs)

[Submitted on 3 Feb 2022 (v1), last revised 6 May 2022 (this version, v2)]

Title:Matching of Markov Databases Under Random Column Repetitions

Authors:Serhat Bakirtas, Elza Erkip

View PDF

Abstract:Matching entries of correlated shuffled databases have practical applications ranging from privacy to biology. In this paper, motivated by synchronization errors in the sampling of time-indexed databases, matching of random databases under random column repetitions and deletions is investigated. It is assumed that for each entry (row) in the database, the attributes (columns) are correlated, which is modeled as a Markov process. Column histograms are proposed as a permutation-invariant feature to detect the repetition pattern, whose asymptotic-uniqueness is proved using information-theoretic tools. Repetition detection is then followed by a typicality-based row matching scheme. Considering this overall scheme, sufficient conditions for successful matching of databases in terms of the database growth rate are derived. A modified version of Fano's inequality leads to a tight necessary condition for successful matching, establishing the matching capacity under column repetitions. This capacity is equal to the erasure bound, which assumes the repetition locations are known a-priori. Overall, our results provide insights on privacy-preserving publication of anonymized time-indexed data.

Subjects:	Information Theory (cs.IT)
Cite as:	arXiv:2202.01730 [cs.IT]
	(or arXiv:2202.01730v2 [cs.IT] for this version)
	https://doi.org/10.48550/arXiv.2202.01730

Submission history

From: Serhat Bakirtas [view email]
[v1] Thu, 3 Feb 2022 17:48:04 UTC (320 KB)
[v2] Fri, 6 May 2022 22:38:15 UTC (112 KB)

Computer Science > Information Theory

Title:Matching of Markov Databases Under Random Column Repetitions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Theory

Title:Matching of Markov Databases Under Random Column Repetitions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators