Distributed and Online Machine Unlearning for Streaming AI Systems
This initiative establishes a joint University of Toronto–University of Warwick collaboration on distributed and online machine unlearning for streaming AI systems. Machine unlearning enables AI systems to remove the influence of specific data without full retraining—an increasingly critical capability for privacy compliance, data deletion requests, poisoned-data removal, and trustworthy AI governance. Existing unlearning research largely assumes static, offline-trained models and does not address modern AI systems that learn continuously over streaming data and execute across distributed infrastructures.
The project will investigate the systems and algorithmic foundations required for scalable and verifiable unlearning in distributed online-learning environments. Core research thrusts include: (1) tracking data-to-model influence in streaming systems; (2) rollback, replay, and corrective mechanisms for selective forgetting; (3) tradeoffs between exact and approximate online unlearning; and (4) consistency and verification mechanisms for distributed unlearning operations.
The collaboration combines the complementary expertise of Professor Hans-Arno Jacobsen in distributed stream/data processing and scalable systems with Professor Peter Triantafillou’s expertise in data management, machine learning systems, and privacy-aware analytics. Expected outcomes include joint publications, prototype architectures, benchmarking methodologies, and follow-on international funding proposals, establishing a sustained strategic partnership in trustworthy AI infrastructure and adaptive data-intensive systems.