In June we ran Oriented R-CNN on a MacBook with --device mps — no CUDA toolchain, one CLI command, rotated boxes on a real aerial tile. v0.2.0 added a fourth detector family: Rotated FCOS, an anchor-free single-stage model with a decoded-rIoU 1× Hub checkpoint.

This post puts both on the same Mac and the same images. Same canvas, same NMS — Apple M1 Max, PyTorch MPS, Hub weights. We also pull the training wall times from the published Hub runs (same NVIDIA L4 recipe). The question is practical: does the new one-stage model feel good enough on a laptop to replace the two-stage demo default?

Short answer: FCOS is the faster one-stage / cheaper-train / MPS demo. It is not more accurate: official Task 1 is 73.07% versus Oriented R-CNN 76.73%. Scores are less peaked, so copying a two-stage threshold onto FCOS will silently drop half the boxes. Use 0.20 for FCOS and 0.55 for Oriented R-CNN.


What we compare Link to heading

Oriented R-CNN 1×Rotated FCOS 1×
Hub slugoriented_rcnn_dota_le90_1xrotated_fcos_dota_le90_1x
Official Task 1 AP5076.73%73.07%
Architecturetwo-stage (RPN + oriented RoIAlign)one-stage, anchor-free
Parameters41.3M36.2M
Checkpoint size315 MB276 MB
1× train wall (NVIDIA L4)1d 12h 22m7h 59m (~4.5× faster)
Mean epoch (incl. periodic mAP)~3h 1m~40m

Both are ResNet-50 + FPN, DOTA le90, train+val pretrain. Published scores are official Task 1, not local val.

Inference numbers below: Apple M1 Max, 64 GB unified memory, PyTorch 2.13.0, oriented-det 0.2.0, --device mps. Protocol: deploy score floors, merge NMS IoU ≤ 0.1. Training wall times come from the Hub run logs on a single NVIDIA L4 (see below).


Quick start Link to heading

From an oriented-det checkout with the usual macOS install (uv pip install torch torchvision then uv pip install -e .):

odet pretrained download rotated_fcos_dota_le90_1x
odet pretrained download oriented_rcnn_dota_le90_1x

odet image-demo demo/demo.jpg hf://rotated_fcos_dota_le90_1x \
  --out-file demo_fcos.png \
  --device mps \
  --score-thr 0.20 \
  --nms-thr 0.1

Swap the slug for Oriented R-CNN and use --score-thr 0.55. Keep --nms-thr 0.1.


Bus lot: same scene as the June demo Link to heading

demo/demo.jpg is the 1024×1024 DOTA tile from the macOS walkthrough — diagonal buses and trucks, the scene where axis-aligned boxes look silly.

Input: demo.jpg — DOTA aerial bus lot

Input: demo.jpg — DOTA aerial bus lot

Oriented R-CNN 1× — demo.jpg at score ≥ 0.55

Oriented R-CNN 1× — demo.jpg at score ≥ 0.55

Rotated FCOS 1× — demo.jpg at score ≥ 0.20

Rotated FCOS 1× — demo.jpg at score ≥ 0.20

Box geometry looks right on both: headings follow the chevron parking, large-vehicle / small-vehicle labels match. The visible difference is score calibration — Oriented R-CNN piles many boxes near 1.00; FCOS spreads them across a wider band. That is expected for a one-stage sigmoid head versus a two-stage ROI classifier, but it changes which CLI threshold you want.


Score thresholds: do not copy a two-stage --score-thr onto FCOS Link to heading

The June post uses --score-thr 0.55 for Oriented R-CNN (the Hub deploy floor). A stricter overlay such as 0.7 still works on that family. On FCOS it does not: a two-stage threshold hides most of the scene.

DetectorDeploy --score-thrWhy
Rotated FCOS 1×0.20sigmoid head; F1 peaks near 0.25 on the local val sweep
Oriented R-CNN 1×0.55peaked two-stage scores

For FCOS demos, start at 0.20. For Oriented R-CNN, 0.55. Keep NMS at 0.1.


Latency on MPS Link to heading

Timed after warmup; mean of 5 single-forward runs (or 3 for the tiled harbor). Canvas 1024×1024, ORIENTED_DET auto window batch 32 on MPS.

ImageWindowsOriented R-CNNRotated FCOSSpeedup
Sparse DOTA tiles (avg of 4)10.38 s0.20 s~1.9×
demo.jpg (dense vehicles)10.48 s0.42 s~1.15×
large.jpg harbor62.98 s1.73 s~1.7×

On sparse tiles the one-stage head is almost faster. On the dense bus lot the gap shrinks because final oriented NMS (Python, AABB-prefiltered) scales with detection count. The harbor tile (six overlapping windows) still favors FCOS by about 40%.

Neither model needs custom CUDA kernels. MPS just works.


Training wall time (same L4 recipe) Link to heading

The Hub checkpoints were trained on the same DOTA tile recipe: 13,691 train+val tiles, batch size 2, 12 epochs, single NVIDIA L4. Timings are from the published sidecar logs (oriented_rcnn_r50_fpn_dota_le90_1x-725c244f.log, rotated_fcos_r50_fpn_dota_le90_1x-a87b6dba.log).

ModelTotal wallMean epoch
Oriented R-CNN 1×1d 12h 22m~3h 1m
Rotated FCOS 1×7h 59m~40m

Rotated FCOS finishes the 1× schedule in an afternoon — about 4.5× less wall clock than Oriented R-CNN on the same GPU and data.

That gap is architectural: Oriented R-CNN pays for an RPN plus oriented RoIAlign on every proposal every step. FCOS is a single dense head over P3–P7. The July ProbIoU post already showed Rotated Faster R-CNN (~58 min/epoch, 11h 30m 1× wall) beating Oriented R-CNN on training cost; FCOS lands in a similar per-epoch band while staying one-stage and anchor-free.


Harbor: sliding-window ships Link to heading

Same large.jpg as the v0.1.1 harbor demo (1299×1904). The DOTA recipe tiles oversized rasters; here that is six 1024 windows with 200 px overlap.

Input: large.jpg — marina with ships at many headings

Input: large.jpg — marina with ships at many headings

Oriented R-CNN 1× — large.jpg at score ≥ 0.55

Oriented R-CNN 1× — large.jpg at score ≥ 0.55

Rotated FCOS 1× — large.jpg at score ≥ 0.20

Rotated FCOS 1× — large.jpg at score ≥ 0.20

odet image-demo demo/large.jpg hf://rotated_fcos_dota_le90_1x \
  --out-file large_fcos.png \
  --device mps \
  --score-thr 0.20 \
  --nms-thr 0.1

Visually both cover the moored rows; FCOS finishes in under two seconds on this machine.


Planes and storage tanks Link to heading

Two more DOTA val tiles for class variety — airport apron and tank farm — same deploy thresholds.

Planes — Oriented R-CNN 1×

Planes — Oriented R-CNN 1×

Planes — Rotated FCOS 1×

Planes — Rotated FCOS 1×

Both find the aircraft. FCOS scores sit lower; Oriented R-CNN saturates near 1.00. Parking-lot vehicles are comparable.

Storage tanks — Oriented R-CNN 1×

Storage tanks — Oriented R-CNN 1×

Storage tanks — Rotated FCOS 1×

Storage tanks — Rotated FCOS 1×

On official Task 1, FCOS is not ahead of Faster R-CNN on tanks (84.28 vs 84.41). On this tile both FCOS and Oriented R-CNN land the visible tanks cleanly.


When to pick which Link to heading

GoalPick
Fast macOS / MPS demo, one-stage stackrotated_fcos_dota_le90_1x
Fast 1× training iteration on a single L4rotated_fcos_dota_le90_1x (~8 h vs ~1.5 days)
Highest official Task 1 accuracyoriented_rcnn_dota_le90_1x (76.73%)
Throughput / finetune defaultrotated_faster_rcnn_dota_le90_1x (74.42%)
Rotated RoIAlign behaviour / continuity with June tutorialsoriented_rcnn_dota_le90_1x
Ships / elongated classes as the productPrefer Faster R-CNN ProbIoU or Oriented R-CNN; FCOS trails on GTF (59.88 vs 71.39 / 74.61)

For laptop demos after v0.2, Rotated FCOS 1× is the better speed default than Oriented R-CNN: fewer parameters, faster MPS inference, cheaper 1× training, and no RPN. Oriented R-CNN still wins Task 1. Keep FCOS at 0.20 and Oriented R-CNN at 0.55, NMS 0.1.


Commands (copy-paste) Link to heading

# Prefetch
odet pretrained download rotated_fcos_dota_le90_1x
odet pretrained download oriented_rcnn_dota_le90_1x

# FCOS on the June bus-lot tile
odet image-demo demo/demo.jpg hf://rotated_fcos_dota_le90_1x \
  --out-file demo_fcos.png --device mps --score-thr 0.20 --nms-thr 0.1

# Side-by-side Oriented R-CNN
odet image-demo demo/demo.jpg hf://oriented_rcnn_dota_le90_1x \
  --out-file demo_orcnn.png --device mps --score-thr 0.55 --nms-thr 0.1

# Harbor (sliding windows)
odet image-demo demo/large.jpg hf://rotated_fcos_dota_le90_1x \
  --out-file large_fcos.png --device mps --score-thr 0.20 --nms-thr 0.1

References Link to heading


Written on September 2, 2026 by Jeff Faudi. Link to heading