Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
The research identifies a privacy vulnerability in document-processing multimodal LLMs where the models leak memorized training data relationships to infer missing information when visual evidence is insufficient. This reveals domain-specific privacy risks in models trained on sensitive identity documents that differ from general-purpose MLLM concerns.
Why this matters
The research identifies a privacy vulnerability in document-processing multimodal LLMs where the models leak memorized training data relationships to infer missing information when visual evidence is insufficient. This reveals domain-specific privacy risks in models trained on sensitive identity documents that differ from general-purpose MLLM concerns.
Check the original work
This explanation is Korpalis’s guide to the material, not a replacement for it. Read the publisher’s page for the full method, evidence and limitations.