Understand video with the local models — summarize, caption, and index clips on your box
August 2026 · SearchAI Inference Server team
videosummarize + caption Qwen visionexact API calls publishing workflows
The same vision models that read a still image also read video — frames sampled over time, so the model sees a sequence, not one picture. That's the piece a publishing pipeline needs: turn a screen recording, a product demo, or a tutorial into a summary, a caption, chapter steps, or an index entry — automatically, and entirely on your own box. No upload, no per-minute video API. This post shows it working on two sample clips, with the exact calls to copy. Both clips are synthetic, generated for this article.
The exact call
Video rides the normal chat endpoint — one content part of type
video_url carrying the clip as a data URL. Base64-encode the
file and send it:
B64=$(base64 -w0 clip.mp4)
curl http://<host>:8081/v1/chat/completions \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{
"model": "q35-4b",
"messages": [{"role":"user","content":[
{"type":"text","text":"List the steps shown in this video, in order."},
{"type":"video_url","video_url":{"url":"data:video/mp4;base64,'"$B64"'"}}
]}]
}'
The server samples frames with ffmpeg and feeds them to the model — you
just send the file. That's the whole integration: it's the chat API you
already use, with a video_url part instead of an
image_url one.
Example 1 — a how-to clip → steps and a caption
Ask for the steps ("List the steps shown in this video, in
order.") and the 4B returns them cleanly:
1. Login
2. Submit Claim
Swap the prompt for a caption
("Write a single one-sentence caption for this video, suitable for
a help-center listing.") and you get publish-ready metadata:
This video demonstrates the login and submission process for
creating a new identity in the system.
Same clip, same endpoint — the prompt decides whether you get chapter steps for a tutorial index or a one-line caption for a listing.
Example 2 — a status animation → a summary
The model doesn't just read one frame — it follows the sequence.
Prompted with "Summarize what this tracking video shows in two
sentences.":
This tracking video shows the progress of shipment TRK-556677 through
its delivery journey, starting from being received at a facility and
moving along a timeline. The package advances through stages including
sorted, out for delivery, and finally delivered, with each status
update marked by the package's position on the timeline.
It caught the tracking number, the ordered stages (received → sorted → out for delivery → delivered), and the fact that it's a timeline — from the frames alone.
What this unlocks for publishing
Point the same call at your own library and the output drops straight into a content pipeline:
- Auto-summaries & descriptions — generate the blurb for every clip in a catalog instead of writing them by hand.
- Chapters & step lists — turn tutorials and screencasts into navigable, searchable steps.
- Captions & titles — publish-ready one-liners for listings and thumbnails.
- Search index & tags — extract entities, tags, and keywords so your video is findable as text.
- Review & QA — flag whether a clip actually shows the steps it claims, before it ships.
Because it runs locally, you can batch a whole archive without a per-minute bill, and unreleased footage never leaves your network — which matters when the video is a product not yet announced.
The two sample clips are yours to try — run the exact call above, then point it at your own footage:
Install the server (Linux, one line) — video
understanding comes with any vision model (2b, 4b,
9b):
curl -fsSL https://inference-server.searchblox.com/install | sudo bash
Video outputs above were captured from q35-4b
on a running SearchAI Inference Server, August 2026, from the synthetic
clips shown (ffmpeg handles frame sampling; the host needs ffmpeg, which
the installer sets up). See the model catalog
for the vision line-up and
the image-extraction post
for the still-image companion.