From 9e7855ffdd38fdd7ab8e2753e0e3dcb0240372fd Mon Sep 17 00:00:00 2001 From: Harshini Pulivarthi Date: Thu, 30 Jul 2026 12:52:27 -0700 Subject: [PATCH 1/3] Changed exisiting slalom file structure to match: classical, yolo, tests --- perception/tasks/slalom/.gitignore | 10 + perception/tasks/slalom/README.md | 128 ++++ perception/tasks/slalom/__init__.py | 0 perception/tasks/slalom/classical/__init__.py | 17 + perception/tasks/slalom/classical/detector.py | 631 ++++++++++++++++++ .../tasks/slalom/slalom_sequence_tracker.py | 141 ++++ .../tasks/slalom/tests/test_detector.py | 111 +++ perception/tasks/slalom/yolo/README.md | 5 + 8 files changed, 1043 insertions(+) create mode 100644 perception/tasks/slalom/.gitignore create mode 100644 perception/tasks/slalom/README.md create mode 100644 perception/tasks/slalom/__init__.py create mode 100644 perception/tasks/slalom/classical/__init__.py create mode 100644 perception/tasks/slalom/classical/detector.py create mode 100644 perception/tasks/slalom/slalom_sequence_tracker.py create mode 100644 perception/tasks/slalom/tests/test_detector.py create mode 100644 perception/tasks/slalom/yolo/README.md diff --git a/perception/tasks/slalom/.gitignore b/perception/tasks/slalom/.gitignore new file mode 100644 index 0000000..1101f16 --- /dev/null +++ b/perception/tasks/slalom/.gitignore @@ -0,0 +1,10 @@ +.DS_Store +.venv/ +GOPR1146 copy.MP4 +outputs/ +train_images +__pycache__/ +*.pyc +*.egg-info/ +test_images +.pytest_cache \ No newline at end of file diff --git a/perception/tasks/slalom/README.md b/perception/tasks/slalom/README.md new file mode 100644 index 0000000..94b2a31 --- /dev/null +++ b/perception/tasks/slalom/README.md @@ -0,0 +1,128 @@ +# urb-slalom + +Prototype perception algorithms for RoboSub Task 2: Avoid Debris / Slalom. + +The task has three repeated pipe sets arranged as: + +```text +WHITE RED WHITE +``` + +The AUV should navigate through each set while keeping the red pipe on the same side it used when passing the gate, and while staying vertically within the pipe area. + +## What This Repo Contains + +- A classical OpenCV detector for red and white vertical PVC pipes. +- A high-level slalom target estimator that converts pipe detections into a steering target. +- A CLI for testing the detector on images or video. +- Synthetic unit tests that exercise the core detector without needing pool footage. + +This is meant to be a fast iteration sandbox before porting the algorithm into the ROS 2 `cv` package. + +## Install + +```bash +python3 -m venv .venv +source .venv/bin/activate +python -m pip install -e ".[dev]" +``` + +## Run On An Image + +```bash +urb-slalom path/to/frame.png --show +``` + +Write an annotated result: + +```bash +urb-slalom path/to/frame.png --output annotated.png +``` + +Use the opposite pass side: + +```bash +urb-slalom path/to/frame.png --red-side left +``` + +## Run On A Directory Of Images + +Point the CLI at a folder instead of a single file to browse a whole dataset. It prints an estimate per image, and with `--show` pops up a window, pausing on each frame until you press a key (`q`/Esc quits early). + +```bash +# Live popup, one keypress per image +urb-slalom "Slalom Task/test/images" --show + +# Or write annotated copies to inspect later, no popup needed +urb-slalom "Slalom Task/test/images" --output outputs/annotated_test +``` + +`Slalom Task/` is a labeled Roboflow dataset (`train/`, `valid/`, `test/`, each with `images/` + YOLO-format `labels/`, plus `data.yaml` with classes `red`/`white`) — useful for testing detection against real course footage rather than just synthetic frames. + +## Test Proxy Footage With One White Pole + +If you have course footage that is not true slalom footage, but has a similar vertical white pole, use `white-pole` mode. This mode ignores the full red/white slalom layout and simply places a yellow target dot to the left or right of the detected white pole. + +```bash +mkdir -p outputs +urb-slalom "GOPR1146 copy.MP4" --mode white-pole --target-side left --output outputs/gopro_white_pole_left.mp4 +``` + +Or use the helper script: + +```bash +./scripts/test_gopro_white_pole.sh +``` + +The annotated output draws: + +- a white bounding box around the detected white pole +- a yellow dot on the requested side of the pole +- a yellow line from the bottom-center of the frame to the target dot + +This is only a proxy test. It is useful for tuning white-pole segmentation and target placement, but it does not validate the full slalom behavior because true slalom needs red and white pipe-set reasoning. + +## Run Tests + +```bash +pytest +``` + +## Algorithm + +1. Convert the frame to HSV. +2. Segment red pipes with two hue ranges, because red wraps around HSV hue 0. +3. Segment white pipes using low saturation and high value. +4. Clean masks with morphological open/close operations. +5. Extract contours and keep tall, vertical, pipe-like bounding boxes. +6. Pick the best red pipe and the nearest white pipe on each side. +7. Compute a target point in the open corridor that keeps the red pipe on the requested side. + +The detector returns both raw pipe detections and a `SlalomEstimate`: + +- `target_x`: normalized horizontal steering target, where `0.0` is left and `1.0` is right. +- `target_y`: normalized vertical target, useful for staying within the pipe height. +- `yaw_error`: normalized horizontal error from image center. +- `confidence`: simple confidence score based on whether the red pipe and adjacent white pipe were detected. + +### Known Limitation: Color Cast + +Pixel sampling against both the `Slalom Task` dataset and our own GoPro footage shows that at typical course range, water attenuation shifts the red pipe's hue to nearly match the white pipe's and the background's (within a few degrees of hue) — only brightness (HSV Value) reliably separates them. The current hue-based `DetectorConfig` thresholds do not account for this and should be expected to under-detect on real footage until segmentation is reworked around relative brightness rather than hue. + +## ROS 2 Porting Notes + +The next step is to wrap `SlalomDetector` in a ROS 2 node that: + +- subscribes to `/camera/usb/front/compressed` +- publishes debug masks/images +- publishes raw `CVObject` pipe detections +- publishes one high-level `CVObject` target for task planning + +Suggested topics: + +```text +/cv/front/slalom_red_pipe/bounding_box +/cv/front/slalom_white_left/bounding_box +/cv/front/slalom_white_right/bounding_box +/cv/front/slalom_target/bounding_box +``` \ No newline at end of file diff --git a/perception/tasks/slalom/__init__.py b/perception/tasks/slalom/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/perception/tasks/slalom/classical/__init__.py b/perception/tasks/slalom/classical/__init__.py new file mode 100644 index 0000000..8a19363 --- /dev/null +++ b/perception/tasks/slalom/classical/__init__.py @@ -0,0 +1,17 @@ +from perception.tasks.slalom.classical.detector import ( + DetectorConfig, + PassSide, + PipeDetection, + SlalomDetector, + SlalomEstimate, + WhitePoleEstimate, +) + +__all__ = [ + "DetectorConfig", + "PassSide", + "PipeDetection", + "SlalomDetector", + "SlalomEstimate", + "WhitePoleEstimate", +] diff --git a/perception/tasks/slalom/classical/detector.py b/perception/tasks/slalom/classical/detector.py new file mode 100644 index 0000000..a3211b4 --- /dev/null +++ b/perception/tasks/slalom/classical/detector.py @@ -0,0 +1,631 @@ +from __future__ import annotations + +from dataclasses import dataclass +from enum import Enum + +import cv2 +import numpy as np + + +class PassSide(str, Enum): + """Which side of the red pipe the AUV should pass on.""" + + LEFT = "left" + RIGHT = "right" + + +@dataclass(frozen=True) +class PipeDetection: + """A vertical pipe candidate in image coordinates.""" + + label: str + x: int + y: int + width: int + height: int + area: float + score: float + aspect_ratio: float = 0.0 + fill_ratio: float = 0.0 + width_stability: float = 0.0 + vertical_alignment: float = 0.0 + contrast_strength: float = 0.0 + + @property + def center_x(self) -> float: + return self.x + self.width / 2 + + @property + def center_y(self) -> float: + return self.y + self.height / 2 + + @property + def bottom_y(self) -> int: + return self.y + self.height + + +@dataclass(frozen=True) +class SlalomEstimate: + """High-level target for task planning or visual servoing.""" + + target_x: float + target_y: float + yaw_error: float + confidence: float + red_pipe: PipeDetection | None + left_white_pipe: PipeDetection | None + right_white_pipe: PipeDetection | None + + @property + def valid(self) -> bool: + return self.red_pipe is not None and self.confidence > 0 + + +@dataclass(frozen=True) +class WhitePoleEstimate: + """Target estimate for proxy footage with one white pole.""" + + target_x: float + target_y: float + yaw_error: float + confidence: float + white_pipe: PipeDetection | None + target_side: PassSide | None = None + red_pipe: PipeDetection | None = None + + @property + def valid(self) -> bool: + return self.white_pipe is not None and self.confidence > 0 + + +@dataclass(frozen=True) +class DetectorConfig: + red_low_1: tuple[int, int, int] = (0, 60, 40) + red_high_1: tuple[int, int, int] = (15, 255, 255) + red_low_2: tuple[int, int, int] = (160, 60, 40) + red_high_2: tuple[int, int, int] = (179, 255, 255) + white_low: tuple[int, int, int] = (0, 0, 120) + white_high: tuple[int, int, int] = (179, 90, 255) + min_area: int = 120 + min_height_ratio: float = 0.16 + min_aspect_ratio: float = 2.4 + max_aspect_width_ratio: float = 0.22 + morph_kernel_size: int = 5 + + # Contrast-based segmentation (see segment_contrast). At range underwater, + # red pigment is heavily attenuated, so a "red" pole reads as near-black + # rather than red-hued and is indistinguishable from a shadowed white pole + # by hue alone. Poles instead show up as vertical strips that are locally + # much darker (red pole) or much brighter (white pole) than the + # surrounding water, so contrast against a local background estimate + # stands in for color. + contrast_bg_kernel: int = 61 + dark_diff_thresh: int = 18 + bright_diff_thresh: int = 18 + # Sun-rippled water surface also reads as locally dark, producing wide, + # blobby contours that pass the generic pole-shape filter. A real pole is + # much narrower relative to its height than a patch of glare (observed + # aspect ratio ~9-12 for real poles vs. ~2.5-6 for glare blobs in test + # footage), so dark/red contrast candidates need a stricter aspect ratio + # than the hue-based default. + dark_min_aspect_ratio: float = 7.0 + red_pair_min_height_ratio: float = 0.22 + # Candidate-shape classification. These features are intentionally simple + # and inspectable: true PVC pipes are thin, vertically aligned, and keep a + # fairly stable mask width from top to bottom. Glare and shadows may be + # high-contrast, but they usually bulge, smear, or tilt. + min_width_stability: float = 0.35 + min_vertical_alignment: float = 0.65 + + +class SlalomDetector: + """Detect red/white slalom pipes and estimate the desired steering target.""" + + def __init__(self, config: DetectorConfig | None = None) -> None: + self.config = config or DetectorConfig() + + def detect( + self, + frame_bgr: np.ndarray, + pass_side: PassSide = PassSide.LEFT, + method: str = "hsv", + ) -> SlalomEstimate: + """Detect the red/white pipe triplet and estimate a steering target. + + method="hsv" segments by color threshold. method="contrast" segments by + local brightness contrast instead -- use it when red pigment is too + attenuated by water for hue to separate red from white/background (see + segment_contrast()). + """ + if frame_bgr is None or frame_bgr.size == 0: + raise ValueError("frame_bgr must be a non-empty BGR image") + + height, width = frame_bgr.shape[:2] + red_mask, white_mask = self.segment_contrast(frame_bgr) if method == "contrast" else self.segment(frame_bgr) + red_min_aspect_ratio = self.config.dark_min_aspect_ratio if method == "contrast" else None + + red_pipes = self._find_pipe_candidates( + red_mask, + "red", + width, + height, + min_aspect_ratio=red_min_aspect_ratio, + contrast_frame=frame_bgr, + ) + white_pipes = self._find_pipe_candidates(white_mask, "white", width, height, contrast_frame=frame_bgr) + + red_pipe = self._choose_red_pipe(red_pipes, width) + left_white_pipe, right_white_pipe = self._choose_adjacent_whites(red_pipe, white_pipes) + + return self._estimate_target( + image_width=width, + image_height=height, + pass_side=pass_side, + red_pipe=red_pipe, + left_white_pipe=left_white_pipe, + right_white_pipe=right_white_pipe, + ) + + def detect_white_pole_target( + self, + frame_bgr: np.ndarray, + target_side: PassSide | None = PassSide.LEFT, + offset_ratio: float = 0.18, + method: str = "hsv", + ) -> WhitePoleEstimate: + """Detect one white vertical pole and place a target left or right of it. + + If target_side is None, infer the target side from image position: a + pole left of center gets a right-side target, and a pole right of + center gets a left-side target. With only one visible white pole, this + is a proxy for steering back toward the channel interior. + """ + if frame_bgr is None or frame_bgr.size == 0: + raise ValueError("frame_bgr must be a non-empty BGR image") + + height, width = frame_bgr.shape[:2] + red_mask, white_mask = self.segment_contrast(frame_bgr) if method == "contrast" else self.segment(frame_bgr) + white_pipes = self._find_pipe_candidates(white_mask, "white", width, height, contrast_frame=frame_bgr) + white_pipe = self._choose_white_proxy_pipe(white_pipes, width) + + if white_pipe is None: + return WhitePoleEstimate(0.5, 0.5, 0.0, 0.0, None, target_side) + + red_pipe = None + if method == "contrast": + red_candidates = self._find_pipe_candidates( + red_mask, + "red", + width, + height, + min_aspect_ratio=3.6, + min_width_stability=0.12, + min_vertical_alignment=0.45, + contrast_frame=frame_bgr, + ) + red_pipe = self._choose_paired_red_pipe(red_candidates, white_pipe, width, height) + + if red_pipe is not None: + target_x_pixels = (white_pipe.center_x + red_pipe.center_x) / 2 + top_y = min(white_pipe.y, red_pipe.y) + bottom_y = max(white_pipe.bottom_y, red_pipe.bottom_y) + target_x = float(np.clip(target_x_pixels / width, 0.05, 0.95)) + target_y = float(np.clip(((top_y + bottom_y) / 2) / height, 0.05, 0.95)) + confidence = float(np.clip(0.5 * white_pipe.score + 0.5 * red_pipe.score, 0.0, 1.0)) + return WhitePoleEstimate( + target_x=target_x, + target_y=target_y, + yaw_error=target_x - 0.5, + confidence=confidence, + white_pipe=white_pipe, + target_side=target_side, + red_pipe=red_pipe, + ) + + if target_side is None: + target_side = PassSide.RIGHT if white_pipe.center_x < width / 2 else PassSide.LEFT + + direction = -1 if target_side == PassSide.LEFT else 1 + # Apparent pole width grows as the AUV gets closer. Let it nudge the + # target offset wider for close poles while keeping sane image bounds. + base_offset = offset_ratio * width + width_scaled_offset = 14.0 * white_pipe.width + offset_pixels = float(np.clip(max(base_offset, width_scaled_offset), 0.12 * width, 0.32 * width)) + target_x_pixels = white_pipe.center_x + direction * offset_pixels + target_x = float(np.clip(target_x_pixels / width, 0.05, 0.95)) + target_y = float(np.clip(white_pipe.center_y / height, 0.05, 0.95)) + + return WhitePoleEstimate( + target_x=target_x, + target_y=target_y, + yaw_error=target_x - 0.5, + confidence=white_pipe.score, + white_pipe=white_pipe, + target_side=target_side, + red_pipe=None, + ) + + def segment_contrast(self, frame_bgr: np.ndarray) -> tuple[np.ndarray, np.ndarray]: + """Segment poles by local brightness contrast against the water background. + + A large-kernel median blur of the value channel stands in for the local + water brightness (which varies with depth, sun glare, and camera angle). + Pixels much darker than that local background are red-pole candidates; + pixels much brighter are white-pole candidates. Returns (dark_mask, bright_mask). + """ + hsv = cv2.cvtColor(frame_bgr, cv2.COLOR_BGR2HSV) + value = hsv[:, :, 2] + + kernel_size = self.config.contrast_bg_kernel | 1 # medianBlur requires odd + background = cv2.medianBlur(value, kernel_size) + diff = value.astype(np.int16) - background.astype(np.int16) + + dark_mask = (diff < -self.config.dark_diff_thresh).astype(np.uint8) * 255 + bright_mask = (diff > self.config.bright_diff_thresh).astype(np.uint8) * 255 + + red_hue_1 = cv2.inRange(hsv, np.array(self.config.red_low_1), np.array(self.config.red_high_1)) + red_hue_2 = cv2.inRange(hsv, np.array(self.config.red_low_2), np.array(self.config.red_high_2)) + red_hue_mask = cv2.bitwise_or(red_hue_1, red_hue_2) + dark_mask = cv2.bitwise_or(dark_mask, red_hue_mask) + + morph_kernel = np.ones((self.config.morph_kernel_size, self.config.morph_kernel_size), np.uint8) + dark_mask = cv2.morphologyEx(dark_mask, cv2.MORPH_OPEN, morph_kernel) + dark_mask = cv2.morphologyEx(dark_mask, cv2.MORPH_CLOSE, morph_kernel) + bright_mask = cv2.morphologyEx(bright_mask, cv2.MORPH_OPEN, morph_kernel) + bright_mask = cv2.morphologyEx(bright_mask, cv2.MORPH_CLOSE, morph_kernel) + + return dark_mask, bright_mask + + def segment(self, frame_bgr: np.ndarray) -> tuple[np.ndarray, np.ndarray]: + hsv = cv2.cvtColor(frame_bgr, cv2.COLOR_BGR2HSV) + + red_1 = cv2.inRange(hsv, np.array(self.config.red_low_1), np.array(self.config.red_high_1)) + red_2 = cv2.inRange(hsv, np.array(self.config.red_low_2), np.array(self.config.red_high_2)) + red_mask = cv2.bitwise_or(red_1, red_2) + + white_mask = cv2.inRange(hsv, np.array(self.config.white_low), np.array(self.config.white_high)) + + kernel = np.ones((self.config.morph_kernel_size, self.config.morph_kernel_size), np.uint8) + red_mask = cv2.morphologyEx(red_mask, cv2.MORPH_OPEN, kernel) + red_mask = cv2.morphologyEx(red_mask, cv2.MORPH_CLOSE, kernel) + white_mask = cv2.morphologyEx(white_mask, cv2.MORPH_OPEN, kernel) + white_mask = cv2.morphologyEx(white_mask, cv2.MORPH_CLOSE, kernel) + + return red_mask, white_mask + + def annotate(self, frame_bgr: np.ndarray, estimate: SlalomEstimate) -> np.ndarray: + annotated = frame_bgr.copy() + height, width = annotated.shape[:2] + + for pipe in [estimate.left_white_pipe, estimate.red_pipe, estimate.right_white_pipe]: + if pipe is None: + continue + color = (0, 0, 255) if pipe.label == "red" else (255, 255, 255) + cv2.rectangle(annotated, (pipe.x, pipe.y), (pipe.x + pipe.width, pipe.y + pipe.height), color, 2) + cv2.putText( + annotated, + f"{pipe.label}:{pipe.score:.2f}", + (pipe.x, max(20, pipe.y - 8)), + cv2.FONT_HERSHEY_SIMPLEX, + 0.5, + color, + 1, + cv2.LINE_AA, + ) + + if estimate.valid: + target = (int(estimate.target_x * width), int(estimate.target_y * height)) + cv2.drawMarker(annotated, target, (0, 255, 255), cv2.MARKER_CROSS, 20, 2) + cv2.line(annotated, (width // 2, height), target, (0, 255, 255), 2) + + return annotated + + def annotate_white_pole(self, frame_bgr: np.ndarray, estimate: WhitePoleEstimate) -> np.ndarray: + annotated = frame_bgr.copy() + height, width = annotated.shape[:2] + + if estimate.white_pipe is not None: + pipe = estimate.white_pipe + cv2.rectangle(annotated, (pipe.x, pipe.y), (pipe.x + pipe.width, pipe.y + pipe.height), (255, 255, 255), 2) + cv2.putText( + annotated, + f"white:{pipe.score:.2f}", + (pipe.x, max(20, pipe.y - 8)), + cv2.FONT_HERSHEY_SIMPLEX, + 0.5, + (255, 255, 255), + 1, + cv2.LINE_AA, + ) + + if estimate.red_pipe is not None: + pipe = estimate.red_pipe + cv2.rectangle(annotated, (pipe.x, pipe.y), (pipe.x + pipe.width, pipe.y + pipe.height), (0, 0, 255), 2) + cv2.putText( + annotated, + f"red/dark:{pipe.score:.2f}", + (pipe.x, max(20, pipe.y - 8)), + cv2.FONT_HERSHEY_SIMPLEX, + 0.5, + (0, 0, 255), + 1, + cv2.LINE_AA, + ) + + if estimate.valid: + target = (int(estimate.target_x * width), int(estimate.target_y * height)) + cv2.circle(annotated, target, 10, (0, 255, 255), -1) + cv2.circle(annotated, target, 14, (0, 0, 0), 2) + cv2.line(annotated, (width // 2, height), target, (0, 255, 255), 2) + + return annotated + + def _find_pipe_candidates( + self, + mask: np.ndarray, + label: str, + image_width: int, + image_height: int, + min_aspect_ratio: float | None = None, + min_width_stability: float | None = None, + min_vertical_alignment: float | None = None, + contrast_frame: np.ndarray | None = None, + ) -> list[PipeDetection]: + min_aspect_ratio = self.config.min_aspect_ratio if min_aspect_ratio is None else min_aspect_ratio + contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) + candidates: list[PipeDetection] = [] + value: np.ndarray | None = None + background: np.ndarray | None = None + if contrast_frame is not None: + hsv = cv2.cvtColor(contrast_frame, cv2.COLOR_BGR2HSV) + value = hsv[:, :, 2] + background = cv2.medianBlur(value, self.config.contrast_bg_kernel | 1) + + for contour in contours: + area = cv2.contourArea(contour) + if area < self.config.min_area: + continue + + x, y, width, height = cv2.boundingRect(contour) + if width <= 0 or height <= 0: + continue + + aspect_ratio = height / width + height_ratio = height / image_height + width_ratio = width / image_width + + if aspect_ratio < min_aspect_ratio: + continue + if height_ratio < self.config.min_height_ratio: + continue + if width_ratio > self.config.max_aspect_width_ratio: + continue + + fill_ratio, width_stability, vertical_alignment = self._shape_features(mask, contour, x, y, width, height) + width_stability_threshold = ( + min_width_stability + if min_width_stability is not None + else self.config.min_width_stability + if label == "red" + else 0.20 + ) + vertical_alignment_threshold = ( + min_vertical_alignment + if min_vertical_alignment is not None + else self.config.min_vertical_alignment + if label == "red" + else 0.55 + ) + if width_stability < width_stability_threshold: + continue + if vertical_alignment < vertical_alignment_threshold: + continue + + contrast_strength = self._contrast_strength(value, background, contour, label) if value is not None else 0.0 + + vertical_score = min(aspect_ratio / (10.0 if label == "red" else 8.0), 1.0) + size_score = min(height_ratio / 0.5, 1.0) + contrast_score = min(contrast_strength / 45.0, 1.0) + if label == "red": + score = ( + 0.40 * vertical_score + + 0.22 * width_stability + + 0.18 * vertical_alignment + + 0.12 * contrast_score + + 0.08 * size_score + ) + else: + score = ( + 0.30 * vertical_score + + 0.20 * width_stability + + 0.20 * vertical_alignment + + 0.18 * contrast_score + + 0.12 * size_score + ) + + candidates.append( + PipeDetection( + label, + x, + y, + width, + height, + float(area), + float(score), + float(aspect_ratio), + float(fill_ratio), + float(width_stability), + float(vertical_alignment), + float(contrast_strength), + ) + ) + + return sorted(candidates, key=lambda pipe: (pipe.score, pipe.area), reverse=True) + + def _shape_features( + self, + mask: np.ndarray, + contour: np.ndarray, + x: int, + y: int, + width: int, + height: int, + ) -> tuple[float, float, float]: + roi = mask[y : y + height, x : x + width] + row_widths = np.count_nonzero(roi, axis=1) + nonzero_rows = row_widths[row_widths > 0] + + if nonzero_rows.size == 0: + return 0.0, 0.0, 0.0 + + fill_ratio = float(np.count_nonzero(roi) / (width * height)) + median_width = float(np.median(nonzero_rows)) + width_spread = float(np.percentile(nonzero_rows, 90) - np.percentile(nonzero_rows, 10)) + width_stability = float(np.clip(1.0 - width_spread / max(median_width, 1.0), 0.0, 1.0)) + + vertical_alignment = self._vertical_alignment(contour) + return fill_ratio, width_stability, vertical_alignment + + def _vertical_alignment(self, contour: np.ndarray) -> float: + if len(contour) < 5: + return 1.0 + + vx, vy, _, _ = cv2.fitLine(contour, cv2.DIST_L2, 0, 0.01, 0.01).flatten() + vertical_component = abs(float(vy)) / max(float(np.hypot(vx, vy)), 1e-6) + return float(np.clip(vertical_component, 0.0, 1.0)) + + def _contrast_strength( + self, + value: np.ndarray | None, + background: np.ndarray | None, + contour: np.ndarray, + label: str, + ) -> float: + if value is None or background is None: + return 0.0 + + contour_mask = np.zeros(value.shape, dtype=np.uint8) + cv2.drawContours(contour_mask, [contour], -1, 255, -1) + selected = contour_mask > 0 + if not np.any(selected): + return 0.0 + + diff = value.astype(np.int16) - background.astype(np.int16) + candidate_diff = diff[selected] + if label == "red": + return float(max(0.0, -np.median(candidate_diff))) + return float(max(0.0, np.median(candidate_diff))) + + def _choose_red_pipe(self, red_pipes: list[PipeDetection], image_width: int) -> PipeDetection | None: + if not red_pipes: + return None + + image_center = image_width / 2 + return max( + red_pipes, + key=lambda pipe: pipe.score - 0.25 * abs(pipe.center_x - image_center) / image_width, + ) + + def _choose_paired_red_pipe( + self, + red_pipes: list[PipeDetection], + white_pipe: PipeDetection, + image_width: int, + image_height: int, + ) -> PipeDetection | None: + if not red_pipes: + return None + + min_separation = 0.08 * image_width + max_separation = 0.70 * image_width + plausible = [ + pipe + for pipe in red_pipes + if min_separation <= abs(pipe.center_x - white_pipe.center_x) <= max_separation + and pipe.height >= self.config.red_pair_min_height_ratio * image_height + ] + if not plausible: + return None + + white_height = max(float(white_pipe.height), 1.0) + return max( + plausible, + key=lambda pipe: ( + 0.45 * min(pipe.height / white_height, 1.6) + + 0.30 * min(pipe.aspect_ratio / 8.0, 1.0) + + 0.20 * min(pipe.contrast_strength / 60.0, 1.0) + + 0.15 * pipe.vertical_alignment + + 0.10 * pipe.score + - 0.15 * abs(pipe.center_x - image_width / 2) / image_width + ), + ) + + def _choose_adjacent_whites( + self, + red_pipe: PipeDetection | None, + white_pipes: list[PipeDetection], + ) -> tuple[PipeDetection | None, PipeDetection | None]: + if red_pipe is None: + return None, None + + left = [pipe for pipe in white_pipes if pipe.center_x < red_pipe.center_x] + right = [pipe for pipe in white_pipes if pipe.center_x > red_pipe.center_x] + + left_pipe = max(left, key=lambda pipe: (pipe.score, pipe.center_x), default=None) + right_pipe = max(right, key=lambda pipe: (pipe.score, -pipe.center_x), default=None) + return left_pipe, right_pipe + + def _choose_white_proxy_pipe( + self, + white_pipes: list[PipeDetection], + image_width: int, + ) -> PipeDetection | None: + if not white_pipes: + return None + + image_center = image_width / 2 + return max( + white_pipes, + key=lambda pipe: pipe.score - 0.15 * abs(pipe.center_x - image_center) / image_width, + ) + + def _estimate_target( + self, + image_width: int, + image_height: int, + pass_side: PassSide, + red_pipe: PipeDetection | None, + left_white_pipe: PipeDetection | None, + right_white_pipe: PipeDetection | None, + ) -> SlalomEstimate: + if red_pipe is None: + return SlalomEstimate(0.5, 0.5, 0.0, 0.0, None, left_white_pipe, right_white_pipe) + + desired_white = left_white_pipe if pass_side == PassSide.LEFT else right_white_pipe + fallback_offset = 0.22 * image_width + + if desired_white is not None: + target_x_pixels = (red_pipe.center_x + desired_white.center_x) / 2 + top_y = min(red_pipe.y, desired_white.y) + bottom_y = max(red_pipe.bottom_y, desired_white.bottom_y) + confidence = 0.9 + else: + direction = -1 if pass_side == PassSide.LEFT else 1 + target_x_pixels = red_pipe.center_x + direction * fallback_offset + top_y = red_pipe.y + bottom_y = red_pipe.bottom_y + confidence = 0.45 + + target_x = float(np.clip(target_x_pixels / image_width, 0.05, 0.95)) + target_y = float(np.clip(((top_y + bottom_y) / 2) / image_height, 0.05, 0.95)) + yaw_error = target_x - 0.5 + + return SlalomEstimate( + target_x=target_x, + target_y=target_y, + yaw_error=yaw_error, + confidence=confidence, + red_pipe=red_pipe, + left_white_pipe=left_white_pipe, + right_white_pipe=right_white_pipe, + ) diff --git a/perception/tasks/slalom/slalom_sequence_tracker.py b/perception/tasks/slalom/slalom_sequence_tracker.py new file mode 100644 index 0000000..920d918 --- /dev/null +++ b/perception/tasks/slalom/slalom_sequence_tracker.py @@ -0,0 +1,141 @@ +""" +SlalomSequenceTracker +-------------------- +Wraps classical.detector.SlalomDetector to add the pieces the RoboSub +slalom task scores that a single-frame detect() call doesn't give you: + + 1. Multi-gate state — which of the 3 gates you're approaching, and + confirmation that you've passed one before advancing to the next. + 2. Side consistency — remembers which side you passed the red pipe on + at gate 1, flags gate 2/3 if they'd put it on the other side. + 3. Distance/depth cue — approximate real-world range to the current gate + from its known 1524mm (60in) white-to-white width vs. the pixel span + your detector already computes between left_white_pipe/right_white_pipe. + 4. Stagger-informed bias — after a gate is passed, the next gate sits + ~1000mm laterally offset (per the course drawing), alternating side. + Useful to bias where you point the camera/search first. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from enum import Enum + +from perception.tasks.slalom.classical.detector import PassSide, SlalomEstimate + +# --- Known course geometry (mm), from the design drawing ------------------- +GATE_WIDTH_MM = 1524.0 # white-to-white pipe center distance +GATE_STAGGER_MM = 1000.0 # lateral offset of middle gate from outer gates +GATE1_TO_GATE3_MM = 2002.9 # longitudinal spacing, gate 1 to gate 3 centerline + + +class GateIndex(Enum): + GATE_1 = 0 + GATE_2 = 1 + GATE_3 = 2 + DONE = 3 + + +@dataclass +class GateResult: + index: GateIndex + pass_side: PassSide | None + distance_mm: float | None + confidence: float + + +@dataclass +class TrackerState: + current_gate: GateIndex = GateIndex.GATE_1 + committed_side: PassSide | None = None + side_mismatches: int = 0 + gates_passed: int = 0 + _was_valid_last_frame: bool = False + _last_red_center_x: float | None = None + + +class SlalomSequenceTracker: + """ + Call update(estimate, image_width) once per frame with the SlalomEstimate + SlalomDetector.detect() already returns. Everything below reuses fields + that already exist on SlalomEstimate/PipeDetection — no changes needed + to detector.py itself. + """ + + def __init__(self, focal_length_px: float): + # From camera calibration (fx). If uncalibrated, derive empirically: + # fx = (pixel_width_at_known_range_mm * range_mm) / GATE_WIDTH_MM + # using a frame where you know the real distance to a gate. + self.focal_length_px = focal_length_px + self.state = TrackerState() + + def estimate_distance_mm(self, estimate: SlalomEstimate) -> float | None: + if estimate.left_white_pipe is None or estimate.right_white_pipe is None: + return None + pixel_span = abs(estimate.right_white_pipe.center_x - estimate.left_white_pipe.center_x) + if pixel_span <= 0: + return None + return (GATE_WIDTH_MM * self.focal_length_px) / pixel_span + + def update(self, estimate: SlalomEstimate, image_width: int) -> GateResult: + distance_mm = self.estimate_distance_mm(estimate) + result = GateResult( + index=self.state.current_gate, + pass_side=None, + distance_mm=distance_mm, + confidence=estimate.confidence, + ) + + # A gate that was tracked (valid) and just stopped being valid means + # it left the frame — count it as passed and infer which side the + # red pipe was on from the last frame it was still visible. + gate_just_exited = ( + self.state._was_valid_last_frame + and not estimate.valid + and self.state._last_red_center_x is not None + ) + + if gate_just_exited and self.state.current_gate != GateIndex.DONE: + side = ( + PassSide.LEFT + if self.state._last_red_center_x < image_width / 2 + else PassSide.RIGHT + ) + result.pass_side = side + + if self.state.committed_side is None: + self.state.committed_side = side + elif side != self.state.committed_side: + self.state.side_mismatches += 1 + + self.state.gates_passed += 1 + self.state.current_gate = GateIndex( + min(self.state.current_gate.value + 1, GateIndex.DONE.value) + ) + + self.state._was_valid_last_frame = estimate.valid + if estimate.valid and estimate.red_pipe is not None: + self.state._last_red_center_x = estimate.red_pipe.center_x + + return result + + def expected_search_bias_px(self, px_per_mm: float) -> float: + """ + Signed pixel offset to bias next-frame search toward the expected + 1000mm stagger, alternating direction gate to gate, pointing back + toward the course centerline relative to the committed pass side. + """ + if self.state.gates_passed == 0 or self.state.committed_side is None: + return 0.0 + direction = 1.0 if self.state.committed_side == PassSide.LEFT else -1.0 + alternate = 1 if self.state.gates_passed % 2 == 1 else -1 + return direction * alternate * GATE_STAGGER_MM * px_per_mm + + @property + def run_summary(self) -> dict: + return { + "gates_passed": self.state.gates_passed, + "committed_side": self.state.committed_side.value if self.state.committed_side else None, + "side_mismatches": self.state.side_mismatches, + "side_consistent": self.state.side_mismatches == 0, + } diff --git a/perception/tasks/slalom/tests/test_detector.py b/perception/tasks/slalom/tests/test_detector.py new file mode 100644 index 0000000..ae71282 --- /dev/null +++ b/perception/tasks/slalom/tests/test_detector.py @@ -0,0 +1,111 @@ +import cv2 +import numpy as np + +from perception.tasks.slalom.classical import PassSide, SlalomDetector + + +def make_frame() -> np.ndarray: + frame = np.zeros((480, 640, 3), dtype=np.uint8) + frame[:] = (30, 70, 65) + + cv2.rectangle(frame, (130, 110), (165, 430), (240, 240, 240), -1) + cv2.rectangle(frame, (300, 90), (335, 430), (0, 0, 230), -1) + cv2.rectangle(frame, (475, 130), (510, 430), (240, 240, 240), -1) + return frame + + +def test_detects_pipe_triplet() -> None: + detector = SlalomDetector() + estimate = detector.detect(make_frame(), pass_side=PassSide.LEFT) + + assert estimate.valid + assert estimate.red_pipe is not None + assert estimate.left_white_pipe is not None + assert estimate.right_white_pipe is not None + assert estimate.confidence > 0.8 + + +def test_left_pass_side_targets_left_corridor() -> None: + detector = SlalomDetector() + estimate = detector.detect(make_frame(), pass_side=PassSide.LEFT) + + assert estimate.target_x < estimate.red_pipe.center_x / 640 + assert 0.25 < estimate.target_x < 0.5 + + +def test_right_pass_side_targets_right_corridor() -> None: + detector = SlalomDetector() + estimate = detector.detect(make_frame(), pass_side=PassSide.RIGHT) + + assert estimate.target_x > estimate.red_pipe.center_x / 640 + assert 0.5 < estimate.target_x < 0.75 + + +def test_missing_white_pipe_uses_red_pipe_fallback() -> None: + frame = np.zeros((480, 640, 3), dtype=np.uint8) + frame[:] = (30, 70, 65) + cv2.rectangle(frame, (300, 90), (335, 430), (0, 0, 230), -1) + + detector = SlalomDetector() + estimate = detector.detect(frame, pass_side=PassSide.LEFT) + + assert estimate.valid + assert estimate.confidence < 0.5 + assert estimate.target_x < estimate.red_pipe.center_x / 640 + + +def test_white_pole_mode_targets_left_of_pole() -> None: + frame = np.zeros((480, 640, 3), dtype=np.uint8) + frame[:] = (30, 70, 65) + cv2.rectangle(frame, (300, 90), (335, 430), (240, 240, 240), -1) + + detector = SlalomDetector() + estimate = detector.detect_white_pole_target(frame, target_side=PassSide.LEFT) + + assert estimate.valid + assert estimate.white_pipe is not None + assert estimate.target_x < estimate.white_pipe.center_x / 640 + + +def test_white_pole_mode_targets_right_of_pole() -> None: + frame = np.zeros((480, 640, 3), dtype=np.uint8) + frame[:] = (30, 70, 65) + cv2.rectangle(frame, (300, 90), (335, 430), (240, 240, 240), -1) + + detector = SlalomDetector() + estimate = detector.detect_white_pole_target(frame, target_side=PassSide.RIGHT) + + assert estimate.valid + assert estimate.white_pipe is not None + assert estimate.target_x > estimate.white_pipe.center_x / 640 + + +def test_white_pole_auto_side_targets_toward_image_center() -> None: + frame = np.zeros((480, 640, 3), dtype=np.uint8) + frame[:] = (30, 70, 65) + cv2.rectangle(frame, (120, 90), (155, 430), (240, 240, 240), -1) + + detector = SlalomDetector() + estimate = detector.detect_white_pole_target(frame, target_side=None) + + assert estimate.valid + assert estimate.white_pipe is not None + assert estimate.target_side == PassSide.RIGHT + assert estimate.target_x > estimate.white_pipe.center_x / 640 + + +def test_white_pole_mode_targets_midpoint_when_dark_red_pair_is_visible() -> None: + frame = np.zeros((480, 640, 3), dtype=np.uint8) + frame[:] = (60, 120, 120) + cv2.rectangle(frame, (170, 120), (185, 390), (245, 245, 245), -1) + cv2.line(frame, (420, 90), (455, 410), (15, 25, 25), 12) + + detector = SlalomDetector() + estimate = detector.detect_white_pole_target(frame, target_side=None, method="contrast") + + assert estimate.valid + assert estimate.white_pipe is not None + assert estimate.red_pipe is not None + + expected_midpoint = (estimate.white_pipe.center_x + estimate.red_pipe.center_x) / 2 / 640 + assert abs(estimate.target_x - expected_midpoint) < 0.02 diff --git a/perception/tasks/slalom/yolo/README.md b/perception/tasks/slalom/yolo/README.md new file mode 100644 index 0000000..fd9f6f3 --- /dev/null +++ b/perception/tasks/slalom/yolo/README.md @@ -0,0 +1,5 @@ +# Slalom YOLO (placeholder) + +No YOLO-based slalom detector exists yet — this directory is a placeholder +for one. The current working detector is the classical CV approach in +[`../classical`](../classical). From 79f19de6c19829717acc550d0b1f243887eda5fd Mon Sep 17 00:00:00 2001 From: harshi-puli Date: Thu, 30 Jul 2026 17:21:46 -0700 Subject: [PATCH 2/3] Fixing whitespaces & lint errors --- perception/tasks/slalom/classical/detector.py | 172 +++++++++++++----- .../tasks/slalom/slalom_sequence_tracker.py | 14 +- .../tasks/slalom/tests/test_detector.py | 9 +- 3 files changed, 145 insertions(+), 50 deletions(-) diff --git a/perception/tasks/slalom/classical/detector.py b/perception/tasks/slalom/classical/detector.py index a3211b4..b8ffc77 100644 --- a/perception/tasks/slalom/classical/detector.py +++ b/perception/tasks/slalom/classical/detector.py @@ -141,8 +141,14 @@ def detect( raise ValueError("frame_bgr must be a non-empty BGR image") height, width = frame_bgr.shape[:2] - red_mask, white_mask = self.segment_contrast(frame_bgr) if method == "contrast" else self.segment(frame_bgr) - red_min_aspect_ratio = self.config.dark_min_aspect_ratio if method == "contrast" else None + red_mask, white_mask = ( + self.segment_contrast(frame_bgr) + if method == "contrast" + else self.segment(frame_bgr) + ) + red_min_aspect_ratio = ( + self.config.dark_min_aspect_ratio if method == "contrast" else None + ) red_pipes = self._find_pipe_candidates( red_mask, @@ -152,10 +158,14 @@ def detect( min_aspect_ratio=red_min_aspect_ratio, contrast_frame=frame_bgr, ) - white_pipes = self._find_pipe_candidates(white_mask, "white", width, height, contrast_frame=frame_bgr) + white_pipes = self._find_pipe_candidates( + white_mask, "white", width, height, contrast_frame=frame_bgr + ) red_pipe = self._choose_red_pipe(red_pipes, width) - left_white_pipe, right_white_pipe = self._choose_adjacent_whites(red_pipe, white_pipes) + left_white_pipe, right_white_pipe = self._choose_adjacent_whites( + red_pipe, white_pipes + ) return self._estimate_target( image_width=width, @@ -184,8 +194,14 @@ def detect_white_pole_target( raise ValueError("frame_bgr must be a non-empty BGR image") height, width = frame_bgr.shape[:2] - red_mask, white_mask = self.segment_contrast(frame_bgr) if method == "contrast" else self.segment(frame_bgr) - white_pipes = self._find_pipe_candidates(white_mask, "white", width, height, contrast_frame=frame_bgr) + red_mask, white_mask = ( + self.segment_contrast(frame_bgr) + if method == "contrast" + else self.segment(frame_bgr) + ) + white_pipes = self._find_pipe_candidates( + white_mask, "white", width, height, contrast_frame=frame_bgr + ) white_pipe = self._choose_white_proxy_pipe(white_pipes, width) if white_pipe is None: @@ -203,7 +219,9 @@ def detect_white_pole_target( min_vertical_alignment=0.45, contrast_frame=frame_bgr, ) - red_pipe = self._choose_paired_red_pipe(red_candidates, white_pipe, width, height) + red_pipe = self._choose_paired_red_pipe( + red_candidates, white_pipe, width, height + ) if red_pipe is not None: target_x_pixels = (white_pipe.center_x + red_pipe.center_x) / 2 @@ -211,7 +229,9 @@ def detect_white_pole_target( bottom_y = max(white_pipe.bottom_y, red_pipe.bottom_y) target_x = float(np.clip(target_x_pixels / width, 0.05, 0.95)) target_y = float(np.clip(((top_y + bottom_y) / 2) / height, 0.05, 0.95)) - confidence = float(np.clip(0.5 * white_pipe.score + 0.5 * red_pipe.score, 0.0, 1.0)) + confidence = float( + np.clip(0.5 * white_pipe.score + 0.5 * red_pipe.score, 0.0, 1.0) + ) return WhitePoleEstimate( target_x=target_x, target_y=target_y, @@ -223,14 +243,18 @@ def detect_white_pole_target( ) if target_side is None: - target_side = PassSide.RIGHT if white_pipe.center_x < width / 2 else PassSide.LEFT + target_side = ( + PassSide.RIGHT if white_pipe.center_x < width / 2 else PassSide.LEFT + ) direction = -1 if target_side == PassSide.LEFT else 1 # Apparent pole width grows as the AUV gets closer. Let it nudge the # target offset wider for close poles while keeping sane image bounds. base_offset = offset_ratio * width width_scaled_offset = 14.0 * white_pipe.width - offset_pixels = float(np.clip(max(base_offset, width_scaled_offset), 0.12 * width, 0.32 * width)) + offset_pixels = float( + np.clip(max(base_offset, width_scaled_offset), 0.12 * width, 0.32 * width) + ) target_x_pixels = white_pipe.center_x + direction * offset_pixels target_x = float(np.clip(target_x_pixels / width, 0.05, 0.95)) target_y = float(np.clip(white_pipe.center_y / height, 0.05, 0.95)) @@ -263,12 +287,18 @@ def segment_contrast(self, frame_bgr: np.ndarray) -> tuple[np.ndarray, np.ndarra dark_mask = (diff < -self.config.dark_diff_thresh).astype(np.uint8) * 255 bright_mask = (diff > self.config.bright_diff_thresh).astype(np.uint8) * 255 - red_hue_1 = cv2.inRange(hsv, np.array(self.config.red_low_1), np.array(self.config.red_high_1)) - red_hue_2 = cv2.inRange(hsv, np.array(self.config.red_low_2), np.array(self.config.red_high_2)) + red_hue_1 = cv2.inRange( + hsv, np.array(self.config.red_low_1), np.array(self.config.red_high_1) + ) + red_hue_2 = cv2.inRange( + hsv, np.array(self.config.red_low_2), np.array(self.config.red_high_2) + ) red_hue_mask = cv2.bitwise_or(red_hue_1, red_hue_2) dark_mask = cv2.bitwise_or(dark_mask, red_hue_mask) - morph_kernel = np.ones((self.config.morph_kernel_size, self.config.morph_kernel_size), np.uint8) + morph_kernel = np.ones( + (self.config.morph_kernel_size, self.config.morph_kernel_size), np.uint8 + ) dark_mask = cv2.morphologyEx(dark_mask, cv2.MORPH_OPEN, morph_kernel) dark_mask = cv2.morphologyEx(dark_mask, cv2.MORPH_CLOSE, morph_kernel) bright_mask = cv2.morphologyEx(bright_mask, cv2.MORPH_OPEN, morph_kernel) @@ -279,13 +309,21 @@ def segment_contrast(self, frame_bgr: np.ndarray) -> tuple[np.ndarray, np.ndarra def segment(self, frame_bgr: np.ndarray) -> tuple[np.ndarray, np.ndarray]: hsv = cv2.cvtColor(frame_bgr, cv2.COLOR_BGR2HSV) - red_1 = cv2.inRange(hsv, np.array(self.config.red_low_1), np.array(self.config.red_high_1)) - red_2 = cv2.inRange(hsv, np.array(self.config.red_low_2), np.array(self.config.red_high_2)) + red_1 = cv2.inRange( + hsv, np.array(self.config.red_low_1), np.array(self.config.red_high_1) + ) + red_2 = cv2.inRange( + hsv, np.array(self.config.red_low_2), np.array(self.config.red_high_2) + ) red_mask = cv2.bitwise_or(red_1, red_2) - white_mask = cv2.inRange(hsv, np.array(self.config.white_low), np.array(self.config.white_high)) + white_mask = cv2.inRange( + hsv, np.array(self.config.white_low), np.array(self.config.white_high) + ) - kernel = np.ones((self.config.morph_kernel_size, self.config.morph_kernel_size), np.uint8) + kernel = np.ones( + (self.config.morph_kernel_size, self.config.morph_kernel_size), np.uint8 + ) red_mask = cv2.morphologyEx(red_mask, cv2.MORPH_OPEN, kernel) red_mask = cv2.morphologyEx(red_mask, cv2.MORPH_CLOSE, kernel) white_mask = cv2.morphologyEx(white_mask, cv2.MORPH_OPEN, kernel) @@ -297,11 +335,21 @@ def annotate(self, frame_bgr: np.ndarray, estimate: SlalomEstimate) -> np.ndarra annotated = frame_bgr.copy() height, width = annotated.shape[:2] - for pipe in [estimate.left_white_pipe, estimate.red_pipe, estimate.right_white_pipe]: + for pipe in [ + estimate.left_white_pipe, + estimate.red_pipe, + estimate.right_white_pipe, + ]: if pipe is None: continue color = (0, 0, 255) if pipe.label == "red" else (255, 255, 255) - cv2.rectangle(annotated, (pipe.x, pipe.y), (pipe.x + pipe.width, pipe.y + pipe.height), color, 2) + cv2.rectangle( + annotated, + (pipe.x, pipe.y), + (pipe.x + pipe.width, pipe.y + pipe.height), + color, + 2, + ) cv2.putText( annotated, f"{pipe.label}:{pipe.score:.2f}", @@ -320,13 +368,21 @@ def annotate(self, frame_bgr: np.ndarray, estimate: SlalomEstimate) -> np.ndarra return annotated - def annotate_white_pole(self, frame_bgr: np.ndarray, estimate: WhitePoleEstimate) -> np.ndarray: + def annotate_white_pole( + self, frame_bgr: np.ndarray, estimate: WhitePoleEstimate + ) -> np.ndarray: annotated = frame_bgr.copy() height, width = annotated.shape[:2] if estimate.white_pipe is not None: pipe = estimate.white_pipe - cv2.rectangle(annotated, (pipe.x, pipe.y), (pipe.x + pipe.width, pipe.y + pipe.height), (255, 255, 255), 2) + cv2.rectangle( + annotated, + (pipe.x, pipe.y), + (pipe.x + pipe.width, pipe.y + pipe.height), + (255, 255, 255), + 2, + ) cv2.putText( annotated, f"white:{pipe.score:.2f}", @@ -340,7 +396,13 @@ def annotate_white_pole(self, frame_bgr: np.ndarray, estimate: WhitePoleEstimate if estimate.red_pipe is not None: pipe = estimate.red_pipe - cv2.rectangle(annotated, (pipe.x, pipe.y), (pipe.x + pipe.width, pipe.y + pipe.height), (0, 0, 255), 2) + cv2.rectangle( + annotated, + (pipe.x, pipe.y), + (pipe.x + pipe.width, pipe.y + pipe.height), + (0, 0, 255), + 2, + ) cv2.putText( annotated, f"red/dark:{pipe.score:.2f}", @@ -371,7 +433,11 @@ def _find_pipe_candidates( min_vertical_alignment: float | None = None, contrast_frame: np.ndarray | None = None, ) -> list[PipeDetection]: - min_aspect_ratio = self.config.min_aspect_ratio if min_aspect_ratio is None else min_aspect_ratio + min_aspect_ratio = ( + self.config.min_aspect_ratio + if min_aspect_ratio is None + else min_aspect_ratio + ) contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) candidates: list[PipeDetection] = [] value: np.ndarray | None = None @@ -401,27 +467,29 @@ def _find_pipe_candidates( if width_ratio > self.config.max_aspect_width_ratio: continue - fill_ratio, width_stability, vertical_alignment = self._shape_features(mask, contour, x, y, width, height) + fill_ratio, width_stability, vertical_alignment = self._shape_features( + mask, contour, x, y, width, height + ) width_stability_threshold = ( min_width_stability if min_width_stability is not None - else self.config.min_width_stability - if label == "red" - else 0.20 + else self.config.min_width_stability if label == "red" else 0.20 ) vertical_alignment_threshold = ( min_vertical_alignment if min_vertical_alignment is not None - else self.config.min_vertical_alignment - if label == "red" - else 0.55 + else self.config.min_vertical_alignment if label == "red" else 0.55 ) if width_stability < width_stability_threshold: continue if vertical_alignment < vertical_alignment_threshold: continue - contrast_strength = self._contrast_strength(value, background, contour, label) if value is not None else 0.0 + contrast_strength = ( + self._contrast_strength(value, background, contour, label) + if value is not None + else 0.0 + ) vertical_score = min(aspect_ratio / (10.0 if label == "red" else 8.0), 1.0) size_score = min(height_ratio / 0.5, 1.0) @@ -460,7 +528,9 @@ def _find_pipe_candidates( ) ) - return sorted(candidates, key=lambda pipe: (pipe.score, pipe.area), reverse=True) + return sorted( + candidates, key=lambda pipe: (pipe.score, pipe.area), reverse=True + ) def _shape_features( self, @@ -480,8 +550,12 @@ def _shape_features( fill_ratio = float(np.count_nonzero(roi) / (width * height)) median_width = float(np.median(nonzero_rows)) - width_spread = float(np.percentile(nonzero_rows, 90) - np.percentile(nonzero_rows, 10)) - width_stability = float(np.clip(1.0 - width_spread / max(median_width, 1.0), 0.0, 1.0)) + width_spread = float( + np.percentile(nonzero_rows, 90) - np.percentile(nonzero_rows, 10) + ) + width_stability = float( + np.clip(1.0 - width_spread / max(median_width, 1.0), 0.0, 1.0) + ) vertical_alignment = self._vertical_alignment(contour) return fill_ratio, width_stability, vertical_alignment @@ -516,14 +590,17 @@ def _contrast_strength( return float(max(0.0, -np.median(candidate_diff))) return float(max(0.0, np.median(candidate_diff))) - def _choose_red_pipe(self, red_pipes: list[PipeDetection], image_width: int) -> PipeDetection | None: + def _choose_red_pipe( + self, red_pipes: list[PipeDetection], image_width: int + ) -> PipeDetection | None: if not red_pipes: return None image_center = image_width / 2 return max( red_pipes, - key=lambda pipe: pipe.score - 0.25 * abs(pipe.center_x - image_center) / image_width, + key=lambda pipe: pipe.score + - 0.25 * abs(pipe.center_x - image_center) / image_width, ) def _choose_paired_red_pipe( @@ -541,7 +618,9 @@ def _choose_paired_red_pipe( plausible = [ pipe for pipe in red_pipes - if min_separation <= abs(pipe.center_x - white_pipe.center_x) <= max_separation + if min_separation + <= abs(pipe.center_x - white_pipe.center_x) + <= max_separation and pipe.height >= self.config.red_pair_min_height_ratio * image_height ] if not plausible: @@ -571,8 +650,12 @@ def _choose_adjacent_whites( left = [pipe for pipe in white_pipes if pipe.center_x < red_pipe.center_x] right = [pipe for pipe in white_pipes if pipe.center_x > red_pipe.center_x] - left_pipe = max(left, key=lambda pipe: (pipe.score, pipe.center_x), default=None) - right_pipe = max(right, key=lambda pipe: (pipe.score, -pipe.center_x), default=None) + left_pipe = max( + left, key=lambda pipe: (pipe.score, pipe.center_x), default=None + ) + right_pipe = max( + right, key=lambda pipe: (pipe.score, -pipe.center_x), default=None + ) return left_pipe, right_pipe def _choose_white_proxy_pipe( @@ -586,7 +669,8 @@ def _choose_white_proxy_pipe( image_center = image_width / 2 return max( white_pipes, - key=lambda pipe: pipe.score - 0.15 * abs(pipe.center_x - image_center) / image_width, + key=lambda pipe: pipe.score + - 0.15 * abs(pipe.center_x - image_center) / image_width, ) def _estimate_target( @@ -599,9 +683,13 @@ def _estimate_target( right_white_pipe: PipeDetection | None, ) -> SlalomEstimate: if red_pipe is None: - return SlalomEstimate(0.5, 0.5, 0.0, 0.0, None, left_white_pipe, right_white_pipe) + return SlalomEstimate( + 0.5, 0.5, 0.0, 0.0, None, left_white_pipe, right_white_pipe + ) - desired_white = left_white_pipe if pass_side == PassSide.LEFT else right_white_pipe + desired_white = ( + left_white_pipe if pass_side == PassSide.LEFT else right_white_pipe + ) fallback_offset = 0.22 * image_width if desired_white is not None: diff --git a/perception/tasks/slalom/slalom_sequence_tracker.py b/perception/tasks/slalom/slalom_sequence_tracker.py index 920d918..8e688c2 100644 --- a/perception/tasks/slalom/slalom_sequence_tracker.py +++ b/perception/tasks/slalom/slalom_sequence_tracker.py @@ -24,9 +24,9 @@ from perception.tasks.slalom.classical.detector import PassSide, SlalomEstimate # --- Known course geometry (mm), from the design drawing ------------------- -GATE_WIDTH_MM = 1524.0 # white-to-white pipe center distance -GATE_STAGGER_MM = 1000.0 # lateral offset of middle gate from outer gates -GATE1_TO_GATE3_MM = 2002.9 # longitudinal spacing, gate 1 to gate 3 centerline +GATE_WIDTH_MM = 1524.0 # white-to-white pipe center distance +GATE_STAGGER_MM = 1000.0 # lateral offset of middle gate from outer gates +GATE1_TO_GATE3_MM = 2002.9 # longitudinal spacing, gate 1 to gate 3 centerline class GateIndex(Enum): @@ -72,7 +72,9 @@ def __init__(self, focal_length_px: float): def estimate_distance_mm(self, estimate: SlalomEstimate) -> float | None: if estimate.left_white_pipe is None or estimate.right_white_pipe is None: return None - pixel_span = abs(estimate.right_white_pipe.center_x - estimate.left_white_pipe.center_x) + pixel_span = abs( + estimate.right_white_pipe.center_x - estimate.left_white_pipe.center_x + ) if pixel_span <= 0: return None return (GATE_WIDTH_MM * self.focal_length_px) / pixel_span @@ -135,7 +137,9 @@ def expected_search_bias_px(self, px_per_mm: float) -> float: def run_summary(self) -> dict: return { "gates_passed": self.state.gates_passed, - "committed_side": self.state.committed_side.value if self.state.committed_side else None, + "committed_side": ( + self.state.committed_side.value if self.state.committed_side else None + ), "side_mismatches": self.state.side_mismatches, "side_consistent": self.state.side_mismatches == 0, } diff --git a/perception/tasks/slalom/tests/test_detector.py b/perception/tasks/slalom/tests/test_detector.py index ae71282..e2a95d3 100644 --- a/perception/tasks/slalom/tests/test_detector.py +++ b/perception/tasks/slalom/tests/test_detector.py @@ -1,6 +1,5 @@ import cv2 import numpy as np - from perception.tasks.slalom.classical import PassSide, SlalomDetector @@ -101,11 +100,15 @@ def test_white_pole_mode_targets_midpoint_when_dark_red_pair_is_visible() -> Non cv2.line(frame, (420, 90), (455, 410), (15, 25, 25), 12) detector = SlalomDetector() - estimate = detector.detect_white_pole_target(frame, target_side=None, method="contrast") + estimate = detector.detect_white_pole_target( + frame, target_side=None, method="contrast" + ) assert estimate.valid assert estimate.white_pipe is not None assert estimate.red_pipe is not None - expected_midpoint = (estimate.white_pipe.center_x + estimate.red_pipe.center_x) / 2 / 640 + expected_midpoint = ( + (estimate.white_pipe.center_x + estimate.red_pipe.center_x) / 2 / 640 + ) assert abs(estimate.target_x - expected_midpoint) < 0.02 From 1ab3d3ef61df5ad30a66c2fa09bee1e368a0415c Mon Sep 17 00:00:00 2001 From: harshi-puli Date: Thu, 30 Jul 2026 18:28:44 -0700 Subject: [PATCH 3/3] Updated Readme with current structure changes --- perception/tasks/slalom/README.md | 87 +++++++++---------------------- 1 file changed, 26 insertions(+), 61 deletions(-) diff --git a/perception/tasks/slalom/README.md b/perception/tasks/slalom/README.md index 94b2a31..d9b4f7b 100644 --- a/perception/tasks/slalom/README.md +++ b/perception/tasks/slalom/README.md @@ -1,6 +1,6 @@ -# urb-slalom +# slalom -Prototype perception algorithms for RoboSub Task 2: Avoid Debris / Slalom. +Perception algorithms for RoboSub Task 2: Avoid Debris / Slalom. The task has three repeated pipe sets arranged as: @@ -10,82 +10,47 @@ WHITE RED WHITE The AUV should navigate through each set while keeping the red pipe on the same side it used when passing the gate, and while staying vertically within the pipe area. -## What This Repo Contains +## Structure -- A classical OpenCV detector for red and white vertical PVC pipes. -- A high-level slalom target estimator that converts pipe detections into a steering target. -- A CLI for testing the detector on images or video. -- Synthetic unit tests that exercise the core detector without needing pool footage. - -This is meant to be a fast iteration sandbox before porting the algorithm into the ROS 2 `cv` package. - -## Install - -```bash -python3 -m venv .venv -source .venv/bin/activate -python -m pip install -e ".[dev]" -``` - -## Run On An Image - -```bash -urb-slalom path/to/frame.png --show -``` - -Write an annotated result: - -```bash -urb-slalom path/to/frame.png --output annotated.png -``` - -Use the opposite pass side: - -```bash -urb-slalom path/to/frame.png --red-side left +```text +slalom/ +├── classical/ # OpenCV detector for red/white vertical PVC pipes (hsv + contrast segmentation) +├── slalom_sequence_tracker.py # Multi-gate state across frames: side consistency, distance, and search bias +├── yolo/ # Placeholder — no YOLO-based detector exists yet +└── tests/ # Synthetic unit tests that exercise the classical detector without pool footage ``` -## Run On A Directory Of Images - -Point the CLI at a folder instead of a single file to browse a whole dataset. It prints an estimate per image, and with `--show` pops up a window, pausing on each frame until you press a key (`q`/Esc quits early). - -```bash -# Live popup, one keypress per image -urb-slalom "Slalom Task/test/images" --show - -# Or write annotated copies to inspect later, no popup needed -urb-slalom "Slalom Task/test/images" --output outputs/annotated_test -``` +`classical/` is ported from the original `urb-slalom` prototype sandbox. That sandbox's CLI and video/GoPro test scripts were not carried over here — this package exposes the detector and tracker as a plain Python API instead. -`Slalom Task/` is a labeled Roboflow dataset (`train/`, `valid/`, `test/`, each with `images/` + YOLO-format `labels/`, plus `data.yaml` with classes `red`/`white`) — useful for testing detection against real course footage rather than just synthetic frames. +## Usage -## Test Proxy Footage With One White Pole +```python +import cv2 -If you have course footage that is not true slalom footage, but has a similar vertical white pole, use `white-pole` mode. This mode ignores the full red/white slalom layout and simply places a yellow target dot to the left or right of the detected white pole. +from perception.tasks.slalom.classical import PassSide, SlalomDetector +from perception.tasks.slalom.slalom_sequence_tracker import SlalomSequenceTracker -```bash -mkdir -p outputs -urb-slalom "GOPR1146 copy.MP4" --mode white-pole --target-side left --output outputs/gopro_white_pole_left.mp4 -``` +frame = cv2.imread("path/to/frame.png") -Or use the helper script: +detector = SlalomDetector() +estimate = detector.detect(frame, pass_side=PassSide.LEFT) +annotated = detector.annotate(frame, estimate) -```bash -./scripts/test_gopro_white_pole.sh +# Optional: track gate-passing state across a video's frames +tracker = SlalomSequenceTracker(focal_length_px=900.0) +gate_result = tracker.update(estimate, image_width=frame.shape[1]) ``` -The annotated output draws: +Use `detector.detect_white_pole_target(...)` / `detector.annotate_white_pole(...)` for proxy footage that has a single white pole instead of a full red/white pipe set. Pass `method="contrast"` to `detect(...)` for footage where water attenuation makes the red pipe read as near-black rather than red-hued (see Known Limitation below). -- a white bounding box around the detected white pole -- a yellow dot on the requested side of the pole -- a yellow line from the bottom-center of the frame to the target dot +## Test Data -This is only a proxy test. It is useful for tuning white-pole segmentation and target placement, but it does not validate the full slalom behavior because true slalom needs red and white pipe-set reasoning. +`Slalom Task/` is a labeled Roboflow dataset (`train/`, `valid/`, `test/`, each with `images/` + YOLO-format `labels/`, plus `data.yaml` with classes `red`/`white`) — useful for testing detection against real course footage rather than just synthetic frames. `crc-photos/` holds additional still frames from pool/tank footage. ## Run Tests ```bash -pytest +pytest perception/tasks/slalom/tests ``` ## Algorithm