The pump-station node from HW01 is bolted to the wall and running. Now it has to report: both its own health (is the box overheating? is the CPU pegged?) and the pump’s sensors (pressure, flow, vibration). The operations center does not want two feeds in two formats; it wants one uniform telemetry stream it can store, chart, and alert on.

In this assignment you build the collector that produces that stream. It reads the DGX Spark’s real device metrics and the pump’s environment sensors, and emits both as records that share one schema. That stream is the exact input HW03 loads into a time-series database and dashboards in Grafana, so the record format you produce here is a contract, not a detail.

You then run the collector twice, once idle and once under load, and analyze the difference.

Full step-by-step instructions, including every command you need to run, are in the README.md of the repository Lumen creates for you. Work from that README. This page is the summary and the requirements.


Objective and Expected Learning Outcomes

By completing this assignment, you will be able to:

  1. Collect device metrics (CPU, GPU, thermal) from a node’s own kernel interfaces.
  2. Define and emit a structured telemetry record (JSON Lines and CSV).
  3. Merge device health and sensor channels into one consistent schema.
  4. Handle real-world messiness: variable core counts, absent GPU fields, and a failed sensor that must record null instead of crashing the stream.
  5. Generate load and read a utilization-under-load time series.
  6. Reason about whether a workload is CPU-bound or GPU-bound from measured data.

The Device

All real captures are taken on the shared course DGX Spark at cs494.evl.uic.edu. SSH in with the key-based login you set up in HW00 and work under your home folder. Give your node a telemetry identity so your readings are distinguishable from other students’.

Real data only. The collector has a simulate mode so you can develop the pure-Python parts on your laptop, but simulated records are stamped "simulated": true and are not accepted. The captures you submit must come from the DGX Spark; grading re-runs there.

Before collecting anything, run the provided device-info script to see what your node actually reports. It prints every fact this assignment samples, next to the file or command it came from. Anything showing as None there will be null in your telemetry too; that is a property of the device, not a bug in your code, and the collector is built to record it as null rather than crash.


The Telemetry Record (the contract)

Every reading, device or pump, uses the same envelope. Only source and metrics differ. Keep this shape exactly; HW03 depends on it.

A device record carries cpu_avg_percent, gpu_percent, gpu_freq_pct, and temp_c. An environment record carries pump_pressure_psi, flow_rate_lpm, vibration_mm_s, ambient_temp_c, and motor_on. Both share timestamp, device_id, source, and metrics.

Per-core CPU detail is not in the JSON Lines stream, since it would clutter the time-series database. It goes only into the CSV, as core0 through coreN columns.


Instructions

Complete the four TODOs in the collector, in order.

Step 1: Read Device Metrics

Use the provided helpers to return a flat dict with the average CPU percentage, the per-core list, GPU utilization, GPU frequency percentage, and temperature. The docstring names the helpers to use and the exact dict shape to return; wiring them together is your job. The device-robustness helpers are provided and complete, so the same code runs unchanged on any supported device.

Step 2: The Pump Environment Sensor

Return a dict of simulated pump readings: pressure, flow rate, vibration, ambient temperature, and motor state. Ranges are in the docstring.

Step 3: Assemble the Records

This is the hard part and the graded core. Merge the two very different sources into two records that share the envelope above. The device record’s metrics are the four scalars only; the environment record’s metrics are the pump reading. If the run is simulated, add a top-level simulated flag to both.

Step 4: Plot Utilization Under Load

Plot average CPU and GPU utilization against elapsed time, with a mean line, labels, legend, grid, and y-limits of zero to one hundred.

Step 5: Capture Idle, Then Under Load

On the DGX Spark, take a quiet baseline capture first. Then, in a second SSH session, start one of the provided stressors and capture again while it runs. Each run writes a JSON Lines file, a CSV, and a plot.

Step 6: Reflection

Fill in the reflection with your platform, core count, idle-versus-load averages, and the CPU-bound versus GPU-bound analysis.


Submission Requirements

To receive credit, you must:

  1. Work on the development branch in your Lumen-provisioned repository.
  2. Commit your completed collector script.
  3. Commit the idle capture as both JSON Lines and CSV.
  4. Commit the under-load capture as JSON Lines, CSV, and the utilization plot.
  5. Ensure every submitted JSON Lines file is a real capture, with no simulated flag.
  6. Commit your frozen requirements file and completed reflection.
  7. Push to development and open a Student PR from development into main.

Do not merge the Student PR. A missing sensor must record null rather than crash the collector. Do not commit .venv.


Evaluation Criteria

Criterion Weight
TODO 1: device metrics collected via the helpers (real CPU/GPU/thermal from the DGX Spark) 20%
TODO 2: pump environment-sensor reading with the required fields 10%
TODO 3: schema merge: both sources share the exact Lab 03 envelope; scalars in JSONL 25%
JSONL + CSV produced for both idle and under-load captures; real (not simulated) 15%
TODO 4: utilization-under-load plot (labels, mean line, [0,100]) 15%
Reflection: platform, core count, idle-vs-load numbers, CPU-bound vs. GPU-bound analysis 15%

Notes on Course-Wide Requirements

  • A reflection file is required for this assignment, including the certification statement in its header.
  • The workflow is unchanged: development branch, Student PR into main, and no merging of your own PR.
  • This assignment builds directly on Lab 03 (Device and Sensor Telemetry Collection) and the CPU/GPU monitoring from Lab 02. The provided helpers are the same ones from those labs; your job is to wire them together into one clean schema. If you are stuck on a step, read its lab part first. The repository README maps each TODO to its section of the lab.
  • The record envelope you build here is the same one HW03 ingests and HW07 ships over the wire, which is why the schema merge is graded on matching it exactly.
  • Your instructor verifies the shared device each semester. If a step fails in a way that looks like the device rather than your code, ask them rather than trying to debug a machine the whole class shares.

Additional Resources