Behavioral Fingerprints for LLM Endpoint Stability and Identity
Jonah Leshin, Manish Shah, Ian Timmis, et al.
This paper introduces Stability Monitor, a system that tracks whether AI language model endpoints behave consistently over time by periodically testing them with fixed prompts and analyzing how their outputs change. Rather than just checking if a service is running (uptime), it detects when a model's actual behavior shifts due to updates, hardware changes, or other modifications. The system successfully identified various types of changes in controlled tests and found that the same model can behave quite differently depending on which company is hosting it.
LLM monitoringmodel stabilitybehavioral consistencyfingerprinting