← Blog · Models · Performance

Understand video with the local models — summarize, caption, and index clips on your box

August 2026 · SearchAI Inference Server team

videosummarize + caption Qwen visionexact API calls publishing workflows

The same vision models that read a still image also read video — frames sampled over time, so the model sees a sequence, not one picture. That's the piece a publishing pipeline needs: turn a screen recording, a product demo, or a tutorial into a summary, a caption, chapter steps, or an index entry — automatically, and entirely on your own box. No upload, no per-minute video API. This post shows it working on two sample clips, with the exact calls to copy. Both clips are synthetic, generated for this article.

The exact call

Video rides the normal chat endpoint — one content part of type video_url carrying the clip as a data URL. Base64-encode the file and send it:

B64=$(base64 -w0 clip.mp4)

curl http://<host>:8081/v1/chat/completions \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{
    "model": "q35-4b",
    "messages": [{"role":"user","content":[
      {"type":"text","text":"List the steps shown in this video, in order."},
      {"type":"video_url","video_url":{"url":"data:video/mp4;base64,'"$B64"'"}}
    ]}]
  }'

The server samples frames with ffmpeg and feeds them to the model — you just send the file. That's the whole integration: it's the chat API you already use, with a video_url part instead of an image_url one.

Example 1 — a how-to clip → steps and a caption

Ask for the steps ("List the steps shown in this video, in order.") and the 4B returns them cleanly:

1. Login
2. Submit Claim

Swap the prompt for a caption ("Write a single one-sentence caption for this video, suitable for a help-center listing.") and you get publish-ready metadata:

This video demonstrates the login and submission process for
creating a new identity in the system.

Same clip, same endpoint — the prompt decides whether you get chapter steps for a tutorial index or a one-line caption for a listing.

Example 2 — a status animation → a summary

The model doesn't just read one frame — it follows the sequence. Prompted with "Summarize what this tracking video shows in two sentences.":

This tracking video shows the progress of shipment TRK-556677 through
its delivery journey, starting from being received at a facility and
moving along a timeline. The package advances through stages including
sorted, out for delivery, and finally delivered, with each status
update marked by the package's position on the timeline.

It caught the tracking number, the ordered stages (received → sorted → out for delivery → delivered), and the fact that it's a timeline — from the frames alone.

What this unlocks for publishing

Point the same call at your own library and the output drops straight into a content pipeline:

  • Auto-summaries & descriptions — generate the blurb for every clip in a catalog instead of writing them by hand.
  • Chapters & step lists — turn tutorials and screencasts into navigable, searchable steps.
  • Captions & titles — publish-ready one-liners for listings and thumbnails.
  • Search index & tags — extract entities, tags, and keywords so your video is findable as text.
  • Review & QA — flag whether a clip actually shows the steps it claims, before it ships.

Because it runs locally, you can batch a whole archive without a per-minute bill, and unreleased footage never leaves your network — which matters when the video is a product not yet announced.

The two sample clips are yours to try — run the exact call above, then point it at your own footage:

Install the server (Linux, one line) — video understanding comes with any vision model (2b, 4b, 9b):

curl -fsSL https://inference-server.searchblox.com/install | sudo bash

Video outputs above were captured from q35-4b on a running SearchAI Inference Server, August 2026, from the synthetic clips shown (ffmpeg handles frame sampling; the host needs ffmpeg, which the installer sets up). See the model catalog for the vision line-up and the image-extraction post for the still-image companion.