ML for streaming data: introduction and recent advances

Heitor Murilo Gomes and Marco Heyden

Introduction

Data streams:

  • Sequence of items, having a timestamp and a temporal order
  • Data stream models aims at being adaptative and improve themselves as the historic data grows
  • Different from time series:
    • Data streams are processed online, in near-real-time
    • The output of a streaming model is a trainable model, not a trained model

Example:

  • Energy demand prediction
  • Malware detection
  • Etc.

Reasons to use data stream:

  • Impossible to store the data: too great volume, etc
  • We should not store it, for example for privacy issues

Challenges of stream learning:

  • limited time to inspect each element
  • difficult to define train/test set
  • feature / concept evolution

Learning cycle for data stream:

  1. Inspect element at most once
  2. Use limimted memory and computation time
  3. Be ready to predict any point, e.g. for testing
  4. Repeat

Evaluation framework

  • Option 1: we measure the error of each element before using it for training
  • Option 2: windowed testing over time that we represent on a plot rather than a single average

CapyMOA: a framework for adaptative ML on stream data

Interoperability:

  • They leverage the existing implementations in MOA framework
  • Interoperability with Python, pytorch, scikit-learn

Coding:

  • Python
  • Or Java
  • Or both

Some algorithms

Concept drifts: Evolving ML:

  • The stream is changing over time, and we would like our models to handle that
  • Detect and react to changes
  • Lean new concepts and forget old ones

Bagging:

  • We cannot do old-fashing sub-sampling because we do not have all data at once
  • New idea (online bagging) : at each time, a data sample is given to a model that we sample with a Poisson law

PLASTIC

  • Decision tree on streams: the challenge is to know when we have enough data to decide a split on a leave
  • They implement a prunning strategy to re-evaluate previous splits based on most recent data
  • The problem is that it temporarily decreases the perf of the tree
  • PLASTIC is a relax-splitting mechanism that avoids the perf decrease when a node is prunned
  • So they re-structure the tree instead of prunning it

Error detection:

  • Use an auto-encoder and measure the mean reconstruction error
  • Detect anomalies on the error stream, which is easier than on the high-dimensional dataset