The pump-station nodes are streaming telemetry from HW02, but a stream nobody stores or watches is useless. The city wants pump data stored, dashboarded, and alerting before a bearing seizes, not a post-mortem after the pump is dead.

In this assignment you stand up the storage-and-visualization layer of the edge pipeline, InfluxDB and Grafana running as containers, and feed it a real 30-minute pump recording. You then build an anomaly detector that flags the three failure events hidden in that recording. Your detector is judged the way a real one would be: against ground truth. Catching two of three failures is a pump that seized.

Full step-by-step instructions, including every command you need to run, are in the README.md of the repository Lumen creates for you. Work from that README. This page is the summary and the requirements.


Objective and Expected Learning Outcomes

By completing this assignment, you will be able to:

  1. Run a multi-container TSDB and Grafana pipeline with Compose.
  2. Write time-series data to InfluxDB using line protocol, and query it.
  3. Provision a Grafana data source and dashboard automatically.
  4. Build a windowed z-score anomaly detector with a normal-only baseline.
  5. Choose and defend a detection threshold against labeled ground truth (per-event detection, detection latency, false positives).
  6. Reason about why different failure signatures (variance burst, slow ramp, level step) need different features.

The Device

The pipeline and detector run in containers. Develop locally, then run the graded stack on the shared course DGX Spark at cs494.evl.uic.edu.

Because the DGX Spark is shared, Docker host ports are shared across everyone. Use only your own two ports, one for InfluxDB and one for Grafana. They are derived from your login UID, so no two students collide and there is no list to look up; the README gives the exact commands. Prefix your Compose project with your account name.


The PumpWatch Dataset

The dataset is a 30-minute pump recording at 50 Hz, 90,000 rows, with vibration, motor temperature proxy, and motor current proxy channels plus ground-truth labels. A manifest accompanies it: the answer sheet, listing three failure events with exact intervals and the recommended normal-only training window before the first failure.

Event Interval Signature
cavitationBursts 360–540 s vibration variance bursts
bearingWearRamp 840–1200 s rising vibration amplitude (ramp)
overheatStep 1440–1620 s temperature level step

Read the manifest before writing the detector. Because the three signatures differ, your window features must include both a mean and a std per channel; the overheat step is a shift in the mean, invisible to a std-only detector.


The Pipeline

The writer replays the PumpWatch recording into InfluxDB, which stores it, and Grafana dashboards it. The detector runs offline on the CSV, like a batch analysis job, and does not depend on the containers. The pipeline is what makes the stream visible; the detector is what makes a failure actionable.


Instructions

Step 1: Bring Up the Pipeline

Complete the writer and Grafana services in the Compose file; the InfluxDB service is already done. Give the writer its InfluxDB environment variables and mount the data directory read-only, then mount the Grafana provisioning directory plus a state volume. Bring the stack up: InfluxDB starts, the writer replays the recording, and Grafana comes up on your derived port. The writer runs once and exits after writing, which is expected.

Step 2: Complete the Writer and the Dashboard

Complete the writer so each CSV row becomes a pump line-protocol record, and test the format offline first with the dry-run mode. Then fix the Motor Temperature panel’s Flux query, which still filters on a placeholder field name, so the panel shows data. Open Grafana and confirm all three panels draw and that you can see the failure events as disturbances.

The writer’s measurement and field names must match the dashboard’s Flux queries.

Step 3: Complete the Detector

Complete four TODOs: per-channel mean and standard deviation features; a normal-only baseline fitted on training windows before the first failure, with every window z-scored and the score taken as the maximum absolute z across features; a threshold in sigma units; and export of the scores CSV and the plot.

Step 4: Tune and Defend the Threshold

The script prints a ground-truth report: which of the three events were detected, the detection latency of each, and how many windows were false positives. Tune the sigma value, and the window size if needed, until you get 3 of 3 events detected with few false positives. Record the numbers in your reflection; a threshold is only correct if it is defended against the manifest, not because it looks about right.

Step 5: Reflection

Fill in the reflection with your threshold, the per-event detection and latency table, the false-positive count, and the analysis: why a std-only feature misses the overheat step, why the ramp is detected latest, and the threshold versus false-positive trade-off.

Step 6 (optional): Live Continuity

Pipe live device telemetry through the same InfluxDB using the provided HW02 collector, as an optional demonstration that the pipeline is source-agnostic. This does not affect the graded PumpWatch detection. The two paths use different InfluxDB measurements, so they coexist without conflict; do not mix their schemas.


Submission Requirements

To receive credit, you must:

  1. Work on the development branch in your Lumen-provisioned repository.
  2. Commit your completed Compose file, writer, and dashboard JSON.
  3. Commit your completed detector script.
  4. Commit the anomaly scores CSV and the detection plot.
  5. Ensure your detector reports 3 of 3 events detected against the manifest.
  6. Commit your completed reflection.
  7. Push to development and open a Student PR from development into main.

Do not merge the Student PR. Use only your own derived ports. Do not commit .venv, and do not put the InfluxDB token in a committed file.


Evaluation Criteria

Criterion Weight
compose.yaml: three-container stack comes up (writer → InfluxDB → Grafana) 20%
Writer: rowToLineProtocol correct; data lands in InfluxDB 15%
Grafana dashboard: data source + all three panels show the pump stream 10%
Detector TODOs 1–2: mean+std features, normal-only baseline, z-score 20%
Detector TODOs 3–4: threshold catches 3/3 events, CSV + plot exported 20%
Reflection: threshold defended with per-event detection, latency, false positives 15%

Notes on Course-Wide Requirements

  • A reflection file is required for this assignment, including the certification statement in its header.
  • The workflow is unchanged: development branch, Student PR into main, and no merging of your own PR.
  • This assignment builds on Lab 04 (Grafana and TSDB) for the pipeline and Lab 07 Part 8 (threshold alerting) for the detection idea. The pipeline commands are the same ones from Lab 04; here they become a graded, reproducible stack. If you are stuck on a step, read its lab part first. The repository README maps each part to its section of the lab.
  • The detector’s threshold is the one part with no single lab recipe. Lab 07 Part 8 teaches the mechanics, but choosing a value you can defend against the manifest’s ground truth is this assignment’s own work.
  • The HW02 collector is provided complete in this repository so your pipeline continuity does not depend on your own HW02.

Additional Resources