Heitor Murilo Gomes and Marco Heyden
ML for streaming data: introduction and recent advances
Heitor Murilo Gomes and Marco Heyden
Introduction
Data streams:
- Sequence of items, having a timestamp and a temporal order
- Data stream models aims at being adaptative and improve themselves as the historic data grows
- Different from time series:
- Data streams are processed online, in near-real-time
- The output of a streaming model is a trainable model, not a trained model
Example:
- Energy demand prediction
- Malware detection
- Etc.
Reasons to use data stream:
- Impossible to store the data: too great volume, etc
- We should not store it, for example for privacy issues
Challenges of stream learning:
- limited time to inspect each element
- difficult to define train/test set
- feature / concept evolution
Learning cycle for data stream:
- Inspect element at most once
- Use limimted memory and computation time
- Be ready to predict any point, e.g. for testing
- Repeat
Evaluation framework
- Option 1: we measure the error of each element before using it for training
- Option 2: windowed testing over time that we represent on a plot rather than a single average
CapyMOA: a framework for adaptative ML on stream data
Interoperability:
- They leverage the existing implementations in MOA framework
- Interoperability with Python, pytorch, scikit-learn
Coding:
- Python
- Or Java
- Or both
Some algorithms
Concept drifts: Evolving ML:
- The stream is changing over time, and we would like our models to handle that
- Detect and react to changes
- Lean new concepts and forget old ones
Bagging:
- We cannot do old-fashing sub-sampling because we do not have all data at once
- New idea (online bagging) : at each time, a data sample is given to a model that we sample with a Poisson law
PLASTIC
- Decision tree on streams: the challenge is to know when we have enough data to decide a split on a leave
- They implement a prunning strategy to re-evaluate previous splits based on most recent data
- The problem is that it temporarily decreases the perf of the tree
- PLASTIC is a relax-splitting mechanism that avoids the perf decrease when a node is prunned
- So they re-structure the tree instead of prunning it
Error detection:
- Use an auto-encoder and measure the mean reconstruction error
- Detect anomalies on the error stream, which is easier than on the high-dimensional dataset