Ops will not deploy on a vibe. Which YOLO model, at what image size, for the traffic fleet? You answer it with evidence: full latency distributions, a fair model and image-size sweep, and a sustained-load and thermal test, then write a recommendation that survives the steady-state data, not just one lucky cold-start run.

The counting node from HW04 works. Now the operations center needs a defensible call before buying a fleet: bigger models see more but run slower, and a config that looks fast for 100 frames can throttle after ten minutes on a hot rooftop. Your job is to benchmark the node honestly and recommend a configuration you can defend against the numbers.

Full step-by-step instructions, including every command you need to run, are in the README.md of the repository Lumen creates for you. Work from that README. This page is the summary and the requirements.


Objective and Expected Learning Outcomes

By completing this assignment, you will be able to:

  1. Compute latency percentiles (p50/p95/p99), not just the mean.
  2. Time GPU inference honestly, with warmup excluded and synchronized.
  3. Sweep model size against image size fairly, one variable at a time.
  4. Measure sustained-load and thermal drift and separate signal from noise.
  5. Explain why a tiny model can underperform on GPU, through launch overhead and unified memory.
  6. Write a defensible, evidence-based deployment recommendation.

The Device

All work runs on the shared course Jetson AGX Thor at athor00.evl.uic.edu, over SSH, inside the instructor’s GPU container using the same base image as HW04. Thermal behavior is why this assignment lives on a Jetson: a desktop-class box does not throttle the way a fanless edge node does. Confirm the GPU is reachable inside the container before you measure anything.

The GPU is shared; a benchmark is only fair when the device is not overloaded by others, so note whether the device was idle and re-run if needed.

The traffic videos are pre-staged on the device and mounted read-only into the container. Confirm the exact shared path with your instructor; the path in the manifest is a placeholder until then.


Instructions

Step 1: Warm Up the Model

Run untimed warmup passes, then synchronize the GPU before any measurement begins.

Step 2: Benchmark One Configuration

Build the timed, GPU-synchronized per-frame loop and return the statistics: p50, p95, p99, mean, frames per second, and mean vehicles detected. The full distribution matters here, not just the mean.

Step 3: Sweep Model Size Against Image Size

Sweep the configurations, changing one variable at a time so the comparison is fair, and produce one result row per configuration.

Step 4: Run the Sustained-Load Test

Run a long test, compare the first N frames against the last N, and return both a drift percentage and a verdict: real thermal throttling or measurement noise. This is judgment, not plumbing. Latency creep could be throttling or the shared GPU getting busier, so note whether the device was idle, re-run if others were active, and decide with a stated threshold.

Step 5: Generate the Report

Produce the comparison table and a first-cut real-time recommendation.

Step 6: Run the Sweep on the GPU and Write the Report

Bring the stack up with the instructor’s base-image tag, which must match the device’s CUDA stack. This writes the results JSON, results CSV, and report text into the logs directory. Then fill in the report template: paste the table, and argue a recommendation against p95, accuracy in vehicles per frame, the GPU-versus-tiny-model surprise, and the sustained numbers rather than the cold-start burst.

Step 7: Reflection

Fill in the reflection with your real percentiles, sweep observations, thermal verdict, and recommendation.


Submission Requirements

To receive credit, you must:

  1. Work on the development branch in your Lumen-provisioned repository.
  2. Commit your completed benchmark harness.
  3. Commit the benchmark results in JSON and CSV form, and the generated report text.
  4. Commit your filled-in benchmark report.
  5. Ensure the submitted numbers come from a GPU run on the device. A CPU run is fine while developing the harness logic.
  6. Commit your completed reflection.
  7. Push to development and open a Student PR from development into main.

Do not merge the Student PR. Do not commit videos, model weights, or .venv.


Evaluation Criteria

Criterion Weight
TODO 1–2: warmup + GPU-synchronized timing; correct p50/p95/p99 + FPS 25%
TODO 3: fair model × image-size sweep (one variable at a time), one row per config 15%
TODO 4: sustained-load test with a defended throttle-vs-noise verdict 15%
TODO 5 + outputs: report table + benchmarkResults.json/.csv from a real GPU run 15%
benchmarkReport.md: recommendation that cites p95 + accuracy + sustained data 20%
Reflection: percentiles, what the GPU-vs-tiny-model test showed on your device, thermal verdict 10%

The recommendation is graded on whether it survives the sustained-load data, not just the cold-start numbers.


Notes on Course-Wide Requirements

  • A reflection file is required for this assignment, including the certification statement in its header.
  • The workflow is unchanged: development branch, Student PR into main, and no merging of your own PR.
  • This assignment pairs with Lab 06 (Benchmarking and Evaluations). The lab analyzed a latency log on the command line; here you build a harness that sweeps configs, measures the tail and the thermal behavior, and produces a report Ops could act on. If you are stuck on a step, read its lab part first.
  • The HW04 vision pipeline is shipped complete in this repository, so this homework never depends on your own HW04. It is the node you evaluate; you do not edit it. The Dockerfile and Compose file are also provided complete, since containerizing was HW04’s graded work.
  • If the smallest model on GPU is barely faster than CPU, that is expected on some devices and not a bug: kernel-launch overhead dominates a tiny workload on a unified-memory device. If your smallest config does show a healthy GPU speedup, that is also a legitimate result. Report what you measured.
  • Your instructor verifies the shared device each semester. If a step fails in a way that looks like the device rather than your code, ask them rather than trying to debug a machine the whole class shares.

Additional Resources