@jawji/orchestrator
v0.1.1
Published
Standalone onboard vision-assisted autonomy for companion computers, starting with a landing-zone safety check.
Maintainers
Readme
jawji-orchestrator
Standalone onboard vision assisted autonomy for a companion computer, starting with a landing zone safety check.
What this is
This package runs on a companion computer (a Jetson or Raspberry Pi, for example), as its own process, with its own direct connection to the flight controller through a local mavsdk_server. It watches vehicle telemetry, and when a configured mode is triggered (the landing zone check triggers on entering LAND mode), it holds the vehicle, captures a camera frame, sends that frame to a vision language model, and decides whether the current position looks safe.
What this is not
This is not a ground control station feature, and it does not require a ground control station to be connected. It has no dependency on any particular GCS. A GCS, or any other operator interface, can observe and confirm what this package is doing by talking to its local HTTP API, but that is entirely optional and the package works correctly with nothing else connected at all.
This package also does not include a vision language model server, and does not include camera capture code for any specific hardware. Both are supplied by the integrator through the adapter interfaces described below.
How the pieces fit together
mavsdk_serverruns on the companion computer, exposing MAVSDK's gRPC interface against the flight controller connection.- This package connects to that
mavsdk_serverover gRPC, using vendored proto definitions from the officialmavlink/MAVSDK-Protorepository, loaded at runtime with@grpc/proto-loader. There is no official MAVSDK Node.js client published to npm, which is why this package generates its own client rather than depending on one. - When the vehicle's flight mode telemetry reports LAND, the orchestrator commands a hold through the MAVSDK action service, then asks the integrator supplied
CameraSourcefor a frame. - That frame, along with a prompt, is sent to the integrator supplied
VlmClient. This package does not care which vision language model answers, only that the response includes astatusfield (safeorunsafe) and, when unsafe, an optional candidate location. - If the verdict is safe, the vehicle resumes its mission automatically.
- If the verdict is unsafe or the vision language model call failed, the orchestrator holds and exposes the verdict on a small local HTTP server, bound to 127.0.0.1 only, by default on port 48500.
- Something else, a GCS, a remote control channel, anything with network access to that local API, posts a confirm or reject decision to
/confirm. The orchestrator then either repositions to the candidate and re-evaluates, or resumes the original landing. - Step 7 repeats up to three times if repositioning to a candidate still does not come back safe. After three attempts the orchestrator reports a no safe alternative found status and holds indefinitely rather than continuing to reposition on its own.
- If nothing responds to a confirm request at all, the vehicle simply continues holding. This package does not invent a timeout based fallback action. The flight controller's own existing failsafes, such as battery or radio control loss, remain the safety net if the vehicle is genuinely unattended.
Architecture
+---------------------------+
| Companion computer |
| (Jetson, Raspberry Pi) |
| |
Flight <--gRPC-->| mavsdk_server |
controller | ^ |
| | gRPC (vendored |
| | MAVSDK-Proto) |
| v |
| jawji-orchestrator |
| | VisionAssistMode(s) |
| | - LandingZoneCheckMode |
| | VehicleAdapter |
| | CameraSource (yours) |
| | VlmClient (yours) |
| | local status/confirm |
| | HTTP API (127.0.0.1) |
+---------------------------+
^
| HTTP, localhost only
| (you decide what, if
| anything, reaches this)
+---------------------------+
| Optional: a GCS, a relay, |
| jawji-agent, an SSH tunnel |
+---------------------------+The only two hard dependencies are mavsdk_server (for vehicle control and telemetry) and an HTTP endpoint implementing VlmClient (for scene assessment). Everything above the local HTTP API line is optional and this package has no awareness of what, if anything, is there.
Relationship with Jawji Agent
None, currently. This is worth stating plainly since jawji-orchestrator and Jawji Agent get mentioned together often enough that it would be easy to assume they are connected. They are two separate, independently running pieces of software. If you deploy both on the same companion computer today, they run side by side, each unaware the other exists.
The intended integration, not yet built, is for Jawji Agent to poll this package's local /status endpoint and relay /confirm decisions through its own existing authenticated REST API, the same pattern it already uses to expose MediaMTX's local status to a connected Jawji desktop app. Until that exists, anything that wants to observe or confirm what jawji-orchestrator is doing needs to reach its local HTTP API some other way, for example an SSH tunnel, or a small relay process you write yourself.
Miril-Drone-2B-1 support
createMirilVlmClient is a built-in VlmClient implementation for Miril-Drone-2B-1, a vision language model fine-tuned for aerial imagery, with response shapes for scene captioning, visual question answering, and operational coordinate pointing (the operational_coordinate_v2 family used by LandingZoneCheckMode).
Miril has no single canonical HTTP protocol of its own. This client targets the OpenAI-compatible /v1/chat/completions shape instead, because that is what every realistic way of serving Miril implements for multimodal chat: llama-server from llama.cpp, vLLM, and SGLang all speak it. The image is sent as a base64 data URL in an image_url content part, alongside a text instruction built from the requested prompt family plus whatever prompt the calling mode supplies.
import { createOrchestrator, createMirilVlmClient, createLandingZoneCheckMode } from '@jawji/orchestrator';
const orchestrator = createOrchestrator({
mavsdkServerAddress: 'localhost:50051',
vlmClient: createMirilVlmClient({
endpoint: 'http://127.0.0.1:8000/v1/chat/completions', // llama-server, vLLM, or SGLang
promptFamily: 'operational_coordinate_v2',
}),
cameraSource: { async captureFrame() { /* ... */ throw new Error('not implemented'); } },
confirmPolicy: 'gated',
statusServerPort: 48500,
modes: [createLandingZoneCheckMode()],
});This does not ship a Miril server. You still need to stand one up yourself, for example llama-server running a GGUF quantization of Miril-Drone-2B-1 on the companion computer, and point endpoint at it. createHttpVlmClient remains available as the plain, protocol-agnostic option for any other vision language model, or any custom HTTP shape you prefer.
GPS-denied navigation (research, not implemented)
This was asked about directly, so it deserves an honest answer rather than either dismissing it or bolting on unreviewed code that would push position estimates into a real flight controller. There is no landmark memory or GPS-denied mode in this package today. Below is the concrete design it would take, written down so it can be reviewed and built deliberately rather than shipped without scrutiny, given that it touches the flight controller's position estimate.
Why this is a different kind of problem than landing-zone-check. LandingZoneCheckMode is a single-shot check: one frame in, one verdict out, no memory across time. Landmark-based re-localization needs the opposite: persistent state built up over the whole flight, a way to match a new frame against everything remembered so far, and a way to feed a corrected position back into the flight controller.
The three pieces, and what already exists to support each of them:
Recording landmarks while GPS is healthy. MAVSDK's telemetry service has
SubscribeHealth, which includes anis_global_position_okfield, confirmed directly against the realtelemetry.proto. A recording step would run on a timer or distance interval while that flag is true, capturing a frame plus the current position and storing them.Matching a new frame against stored landmarks when GPS is lost. A captioning vision language model like Miril is the wrong tool for this specific step. It is built to describe a scene in words, not to produce a precise, viewpoint-stable embedding for place recognition, and comparing generated captions is a poor substitute for comparing actual visual features. The realistic options, roughly in order of how much they would add to this package's current no-native-addon design goal:
- Perceptual image hashing (average hash or difference hash) as a coarse, cheap first pass. Genuinely implementable in pure JavaScript with a library like
jimp(no native addon, same constraint this package already holds itself to for@grpc/grpc-js), but only catches near-duplicate viewpoints, not real viewpoint-invariant matching. Honest baseline, not real visual place recognition. - A dedicated visual place recognition embedding model (for example NetVLAD-style or a ViT-based place recognition model) run locally via an ONNX runtime, with nearest-neighbor search over stored embeddings. Meaningfully more reliable, but reintroduces a native-or-WASM ML runtime dependency this package has otherwise avoided, and needs a specific model choice and real accuracy validation against actual flight footage before it could be trusted.
- Classical feature matching (ORB or SIFT keypoints via OpenCV) as the well-established SLAM/relocalization approach, at the cost of an OpenCV dependency, which is a native addon on most platforms.
- Perceptual image hashing (average hash or difference hash) as a coarse, cheap first pass. Genuinely implementable in pure JavaScript with a library like
Feeding a match back to the flight controller. MAVSDK's
mocapservice exists for exactly this:MocapService.SetVisionPositionEstimate, confirmed directly against the realmocap.proto, takes a body-frame position estimate (PositionBody: x/y/z in metres, not raw latitude/longitude) intended for exactly the case the proto's own comment describes, "navigation without global positioning sources available (e.g. indoors, or when flying under a bridge)." A landmark match would need converting from a stored global position plus the current relative offset into that body-frame form before calling it.
None of this is built. Given it would be capable of altering the flight controller's position estimate, and the accuracy of any of the matching approaches above needs validating against real flight footage before being trusted, the honest recommendation is to design and review this as its own effort, following the same pattern this package's own LandingZoneCheckMode did: a design document, an explicit confirm-gate default, and real testing before anything resembling autonomous behavior.
Confirm gate policy
By default, confirmPolicy is gated, meaning an unsafe or unknown verdict always blocks on step 7 above until an external confirm arrives. Setting confirmPolicy to autonomous skips that wait entirely and acts on the mode's own default decision immediately. This must be set explicitly. The default is gated on purpose, so that using this package does not silently grant a companion computer the ability to make unattended repositioning decisions unless that is a deliberate choice by whoever is integrating it.
Local HTTP API
Bound to 127.0.0.1 only, not exposed to the network by this package. Default port 48500.
GET /status returns the current orchestrator status as JSON, one of:
{ "state": "idle" }
{ "state": "holding" }
{ "state": "awaiting-confirm", "verdict": { "status": "unsafe", "candidate": { "latitudeDeg": 1.23, "longitudeDeg": 4.56, "description": "clearer patch to the north east" } } }
{ "state": "no-safe-alternative" }POST /confirm accepts a JSON body of either:
{ "action": "accept-candidate" }
{ "action": "reject" }Anything with network reachability to the companion computer's localhost, such as a reverse proxy, an agent process, or an SSH tunnel, can be layered in front of this API to expose it more broadly. This package deliberately does not do that itself, since deciding who is allowed to confirm a landing decision is a security and trust boundary question for the integrator, not something this package should assume an answer to.
Usage
import {
createOrchestrator,
createHttpVlmClient,
createLandingZoneCheckMode,
} from '@jawji/orchestrator';
const orchestrator = createOrchestrator({
mavsdkServerAddress: 'localhost:50051',
vlmClient: createHttpVlmClient('http://127.0.0.1:8000/query'),
cameraSource: {
async captureFrame() {
// Return the current frame as a JPEG or PNG Buffer, from whatever
// camera source is available on this companion computer.
throw new Error('not implemented');
},
},
confirmPolicy: 'gated',
statusServerPort: 48500,
modes: [createLandingZoneCheckMode()],
});
await orchestrator.start();Prerequisites
- A running
mavsdk_serverconnected to the flight controller. - A vision language model reachable over HTTP that accepts an image and a prompt and returns JSON with at least a
statusfield. This package does not ship one. A small local server such asllama-serverfromllama.cpp, or a hosted endpoint, both work equally well as long as the response shape matches.
Development
npm install
npm run build
npm testLicense
MIT, see LICENSE.
