HW04 - Vision at the Edge: Counting and Latency Logging
Late Policy
- Assignments submitted late will incur a deduction of 10% per day, including weekends and university holidays, from the maximum possible score. Submissions beyond three days late will result in a score of 0.
You switch domains here, from the pump station to a traffic intersection node. The city’s traffic team wants vehicle counts at an intersection computed on the node, not streamed to the cloud. You are handed a traffic camera feed and the shared course Jetson AGX Thor.
Two things have to be true for this to ship: the counts must be accurate enough, and each inference must be fast enough to keep up with the video. Your job is to build the counting pipeline and measure its latency the honest way, because the first person to over-report FPS is the engineer who forgot that the GPU runs asynchronously.
Full step-by-step instructions, including every command you need to run, are in the README.md of the repository Lumen creates for you. Work from that README. This page is the summary and the requirements.
Objective and Expected Learning Outcomes
By completing this assignment, you will be able to:
- Run a YOLOv8 inference loop and filter detections to a class set.
- Log per-inference latency in a structured JSON-Lines format.
- Exclude warmup and synchronize the GPU so wall-clock timings are honest.
- Match a container tag to the device’s CUDA stack.
- Containerize a GPU inference service with a GPU device reservation and volumes.
- Judge accuracy, latency, and throughput trade-offs for an edge deployment.
The Device
All work runs on the shared course Jetson AGX Thor at athor00.evl.uic.edu, not on the DGX Spark used in HW00 through HW03. This is the hardware family your group project runs on, so you meet it here first. You install your own key on it, the same way HW00 taught: ssh-copy-id <netid>@athor00.evl.uic.edu. Vehicle counting needs the GPU, so you run inside the instructor’s container, which already has PyTorch, CUDA, OpenCV, and Ultralytics. The GPU is shared; if latency spikes, someone else may be running a job, which is a real property of shared edge hardware.
Confirm inference is available before you measure anything. If CUDA reports as unavailable, GPU access is off or the container tag does not match the device’s CUDA stack; fix that first.
The traffic videos are not in the repository. They are pre-staged on the device and mounted read-only into your container. Confirm the exact shared path with your instructor; the path in the manifest is a placeholder until then.
Instructions
Step 1: Reach the Device and the Inference Environment
SSH to the device and confirm that the inference stack imports and that CUDA is reachable from inside the instructor’s container.
Step 2: Resolve the Device and Load the Model
Map the cpu, cuda, and auto options to the string YOLO expects. For the CUDA case, verify a CUDA GPU is actually usable and raise a helpful error if not. Then load and return the YOLO model.
Step 3: Detect Vehicles
Run the detector on one frame and return both the vehicle-class detections and Ultralytics’ per-stage timing dict, so the caller can log the full latency breakdown without a second inference. Keep only the vehicle classes.
Step 4: Process the Video
This is the hard part and where the grade is won or lost. Build the timed loop: untimed warmup first, then per-frame GPU-synchronized timing, per-class counts, and one latency record per frame via the provided record builder. CUDA is asynchronous; if you stop the timer before the GPU finishes, your latency is fiction and your FPS is impossibly high.
Step 5: Containerize for the GPU
Base the Dockerfile on the instructor-supplied base image tag matching the device’s CUDA stack, set the working directory, copy the scripts, and set a sensible default command. Do not reinstall the inference libraries; they are in the base image. Then complete the Compose service with a ${USER}-prefixed image, a GPU device reservation, the shared videos mounted read-only, a writable logs mount, the command, and the device identity in the environment.
Step 6: Run on the GPU and Capture Results
Bring the stack up to write the counts CSV, summary, per-frame latency JSON Lines, timing file, and annotated frames. Sanity-check your timing: the mean end-to-end latency implies a frame rate that should be in the same ballpark as the pipeline FPS in the timing file. A large gap means you did not synchronize the GPU.
Step 7: Reflection
Fill in the reflection with your real environment, counts, and latency numbers, and the warmup and synchronization explanation.
Submission Requirements
To receive credit, you must:
- Work on the
developmentbranch in your Lumen-provisioned repository. - Commit your completed pipeline script, Dockerfile, and Compose file.
- Commit the counts CSV, summary JSON, per-frame latency JSON Lines, and timing JSON.
- Commit a few annotated detection frames.
- Ensure the submitted numbers come from a GPU run on the device. A CPU run is fine while you develop the logic, but the latency and counts you submit must be the device’s.
- Commit your completed reflection.
- Push to
developmentand open a Student PR fromdevelopmentintomain.
Do not merge the Student PR. Do not commit videos, model weights, or .venv.
Evaluation Criteria
| Criterion | Weight |
|---|---|
TODO 1–2: resolveDevice (cpu/cuda/auto → correct string, GPU check) + loadModel
|
15% |
TODO 3: detectVehicles: vehicle-class filtering, correct detection dicts, returns speed
|
20% |
TODO 4: processVideo: warmup excluded and GPU-synchronized honest timing |
25% |
perFrameLatency.jsonl in the Lab 05 envelope; counts CSV + summary.json produced on the GPU |
15% |
Dockerfile + compose: GPU reservation, correct tag, :ro video mount, writable logs/
|
15% |
| Reflection: real numbers + the warmup/synchronization explanation | 10% |
Latency honesty is graded on method, not fixed numbers: warmup excluded and the GPU synchronized before the clock stops.
Notes on Course-Wide Requirements
- A reflection file is required for this assignment, including the certification statement in its header.
- The workflow is unchanged:
developmentbranch, Student PR intomain, and no merging of your own PR. - This assignment pairs with Lab 05 (YOLO, LocateAnything and Latency Logging). The lab taught the mechanics on single sample images; here you apply them to a real video stream, add correct GPU-synchronized timing, and ship it in a container. The container mechanics come from Lab 02 Parts 16 and 17.
- Lab 05 Part 4 is the one to reread if your FPS looks impossible. Warmup and GPU synchronization are exactly where this assignment’s grade is won or lost.
- The provided helpers in the pipeline script are complete; do not modify them. You implement the four functions.
- Your instructor verifies the shared device each semester. If a step fails in a way that looks like the device rather than your code, ask them rather than trying to debug a machine the whole class shares.
Additional Resources
-
Ultralytics YOLO docs - Model loading, prediction, the results and boxes API, and the
speedtiming dict. -
Ultralytics COCO classes - The class ids you filter on.
-
torch.cuda.synchronize - Making the CPU wait for the GPU, for honest timing.
-
OpenCV VideoCapture - Reading frames and frame rate.
