Vidlizer

by arizawan

152 downloads Not rated yet
GitHub

About

Extract structured JSON from video, images, and PDFs using local LLMs (Ollama, LM Studio, oMLX) or via OpenRouter. Runs fully offline.

Explore

- Any input — local video, image (jpg/png/webp/…), PDF, or URL (YouTube, Loom, Vimeo, Twitter)
- 4 providers — Ollama (fully offline), LM Studio (port 1234), oMLX (Apple Silicon, port 8000), OpenRouter (cloud) — auto-detected in that order
- Cross-provider fallback — primary model fails → automatically switches provider (e.g. oMLX → OpenRouter)
- JSON repair — malformed model output is re-sent to the model to fix before skipping; recovers from partial JSON
- Free-model guard — :free OpenRouter models auto-force concurrency=1 to stay within rate limits
- 3 output formats — --format json (default), summary (plain text by phase), markdown (step-per-section doc)
- Usage tracking — --stats shows per-model token + cost breakdown across all runs; get_usage_stats() MCP tool
- Auto transcript — detects audio, transcribes with Apple MLX Whisper (Neural Engine), merges speech into each flow step
- Perceptual dedup — removes near-duplicate frames before sending (saves tokens)
- analyze_moment — --start/--end flags to focus on a time range
- In-memory cache — repeat runs on the same file skip the API call
- Cost guard — aborts if spend exceeds MAX_COST_USD (default $1.00)
- Live progress — Rich streaming indicator shows elapsed time and token count per batch
- MCP server — use from Claude Code, Cursor, Claude Desktop; provider/model locked via env vars; result includes model_used + provider_used
- Auto-install — missing ffmpeg is brew-installed; mlx-whisper bundled in default install (macOS)
- doctor --fix — interactive repair wizard: installs missing ffmpeg/Ollama/LM Studio via Homebrew, re-runs vidlizer setup for .env, upgrades mlx-whisper
- mcp-setup — one-command MCP config wizard: detects vidlizer-mcp, reads .env, writes editor config or shows a claude mcp add-json one-liner
- Mac-native — file picker dialog, Apple MLX transcription, handles macOS Unicode filenames (e.g. "11:26 AM")

---

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Vidlizer
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository


bash uvx vidlizer setup

Ollama (fully offline, no API key):

bash

ollama pull qwen2.5vl:3b # ~3.2 GB, requires 5 GB+ RAM (recommended)
ollama pull qwen2.5vl:7b # ~6.0 GB, requires 10 GB+ RAM (best quality)


LM Studio (GPU-accelerated local inference):

bash


Copy env.sample to .env:

bash

Every successful run appends a record to ~/.cache/vidlizer/usage.jsonl. Test runs (pytest) are excluded.

vidlizer --stats          # print per-model breakdown
vidlizer --clear-stats    # reset log

Example output:

Usage statistics  (/Users/you/.cache/vidlizer/usage.jsonl)

Total runs: 12
Total tokens in: 58,240
Total tokens out: 6,102
Total cost: ~$0.0312
Total steps: 94

Model Provider Runs Tokens in Tokens out Cost USD
gemma-4-E2B-it-MLX-4bit openai 9 45,800 4,900 free
google/gemini-2.5-flash openrouter 3 12,440 1,202 ~$0.0312

Also available as MCP tool: get_usage_stats() → same data as JSON. clear_usage_stats() resets the log.

---

uvx "vidlizer[mcp]" vidlizer-mcp

pipx inject vidlizer mcp

vidlizer mcp-setup

Interactive wizard — detects vidlizer-mcp, reads your .env, and either writes the config to your editor's config file or shows a one-liner to copy-paste. Supports Claude Code, Cursor, Claude Desktop, Windsurf.

No pip install or pipx needed. uvx downloads and runs vidlizer on the fly:

{
  "mcpServers": {
    "vidlizer": {
      "type": "stdio",
      "command": "uvx",
      "args": ["--from", "vidlizer[mcp]", "vidlizer-mcp"],
      "env": {
        "PROVIDER": "ollama",
        "OLLAMA_HOST": "http://localhost:11434",
        "OLLAMA_MODEL": "gemma4:2b"
      }
    }
  }
}

Swap the env block for any provider (OpenRouter, LM Studio, oMLX) — the command/args stay the same.

---

Use the absolute path to the binary — which vidlizer-mcp gives it. All configs below use JSON (works in Claude Code, Cursor, Claude Desktop, Gemini CLI).

For Claude Code via CLI:

claude mcp add-json vidlizer '{"type":"stdio","command":"/path/to/vidlizer-mcp","env":{"PROVIDER":"openrouter","OPENROUTER_API_KEY":"sk-or-v1-...","OPENROUTER_MODEL":"google/gemini-2.5-flash"}}'

---

| Tool | Returns | Tokens out |
|---|---|---|
| analyze_video(path, **opts) | analysis_id + meta | ~100 |
| list_analyses() | all stored analyses (meta only) | ~50/entry |
| get_summary(id, level) | brief/medium/full text summary | ~200–2K |
| get_step(id, step) | single flow step | ~150 |
| get_steps(id, start, end) | step range | scaled |
| get_phase(id, phase) | all steps in named phase | scaled |
| search_analysis(id, query) | steps matching text (query or keyword) | only hits |
| get_transcript(id, start_s, end_s) | transcript slice | scaled |
| get_full_analysis(id) | full flow + transcript | full |
| delete_analysis(id) | confirmation | ~10 |
| get_usage_stats() | per-model token + cost breakdown | ~50/model |
| clear_usage_stats() | reset usage log | ~10 |

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "vidlizer": {
            "vidlizer": {
                "type": "stdio",
                "command": "uvx",
                "args": [
                    "--from",
                    "vidlizer[mcp]",
                    "vidlizer-mcp"
                ],
                "env": {
                    "PROVIDER": "ollama",
                    "OLLAMA_HOST": "http://localhost:11434",
                    "OLLAMA_MODEL": "gemma4:2b"
                }
            }
        }
    }
}

McpServers

{
    "vidlizer": {
        "type": "stdio",
        "command": "uvx",
        "args": [
            "--from",
            "vidlizer[mcp]",
            "vidlizer-mcp"
        ],
        "env": {
            "PROVIDER": "ollama",
            "OLLAMA_HOST": "http://localhost:11434",
            "OLLAMA_MODEL": "gemma4:2b"
        }
    }
}

Point it at a video, image, or PDF. Get structured JSON — scene by scene.

PyPI
License: MIT
Python 3.10+
macOS
CI
Tests
Buy Me a Coffee
Author
Company

demo

</div>

---

vidlizer pulls frames out of any video, image, or PDF using ffmpeg, sends them to a vision LLM, and returns a flow array — one entry per scene. Each entry tells you what happened, who was on screen, what text was visible, and what changed. If the video has audio, it transcribes it with Apple MLX Whisper and merges the speech into each step.

Runs fully local via Ollama or any OpenAI-compatible server (LM Studio, vLLM, oMLX) — no API key, no data leaving your machine. Or connect OpenRouter for cloud models. vidlizer setup detects what you have installed and writes your config in under a minute.

vidlizer demo.mp4
vidlizer "https://youtube.com/watch?v=..."
vidlizer screenshot.png
vidlizer document.pdf

---

✨ Features

- Any input — local video, image (jpg/png/webp/…), PDF, or URL (YouTube, Loom, Vimeo, Twitter)
- 4 providers — Ollama (fully offline), LM Studio (port 1234), oMLX (Apple Silicon, port 8000), OpenRouter (cloud) — auto-detected in that order
- Cross-provider fallback — primary model fails → automatically switches provider (e.g. oMLX → OpenRouter)
- JSON repair — malformed model output is re-sent to the model to fix before skipping; recovers from partial JSON
- Free-model guard — :free OpenRouter models auto-force concurrency=1 to stay within rate limits
- 3 output formats — --format json (default), summary (plain text by phase), markdown (step-per-section doc)
- Usage tracking — --stats shows per-model token + cost breakdown across all runs; get_usage_stats() MCP tool
- Auto transcript — detects audio, transcribes with Apple MLX Whisper (Neural Engine), merges speech into each flow step
- Perceptual dedup — removes near-duplicate frames before sending (saves tokens)
- analyze_moment — --start/--end flags to focus on a time range
- In-memory cache — repeat runs on the same file skip the API call
- Cost guard — aborts if spend exceeds MAX_COST_USD (default $1.00)
- Live progress — Rich streaming indicator shows elapsed time and token count per batch
- MCP server — use from Claude Code, Cursor, Claude Desktop; provider/model locked via env vars; result includes model_used + provider_used
- Auto-install — missing ffmpeg is brew-installed; mlx-whisper bundled in default install (macOS)
- doctor --fix — interactive repair wizard: installs missing ffmpeg/Ollama/LM Studio via Homebrew, re-runs vidlizer setup for .env, upgrades mlx-whisper
- mcp-setup — one-command MCP config wizard: detects vidlizer-mcp, reads .env, writes editor config or shows a claude mcp add-json one-liner
- Mac-native — file picker dialog, Apple MLX transcription, handles macOS Unicode filenames (e.g. "11:26 AM")

---

📦 Requirements

- macOS (Apple Silicon recommended for transcription speed)
- Python 3.10+
- Ollama mode: Ollama installed + a vision model pulled (5 GB+ RAM)
- LM Studio mode: LM Studio 0.3.16+ with a vision model loaded
- Cloud mode: An OpenRouter API key

ffmpeg is installed automatically via Homebrew on first run if missing.

---

🚀 Install

Option 1 — uvx (no install, run directly)

uvx vidlizer setup

Option 2 — pipx (isolated, globally available)

pipx install vidlizer
vidlizer setup    # interactive wizard: detects providers, writes .env

Option 3 — pip / virtualenv

pip install vidlizer
vidlizer setup

Option 4 — from source

git clone https://github.com/arizawan/vidlizer.git
cd vidlizer
python -m venv .venv && source .venv/bin/activate
pip install -e .
vidlizer setup    # or: cp env.sample .env

First-run wizard

vidlizer setup detects all installed providers, lets you pick primary + fallback, and writes a .env for you. It also offers to pull a vision model for Ollama if none is installed.

$ vidlizer setup
  Detected providers:
    1.  Ollama        → qwen2.5vl:3b
    2.  OpenRouter    → google/gemma-3-27b-it:free

Primary provider (1–2): 1
Fallback (1–1, Enter to skip): 2

✓ .env written → /your/project/.env

Health check

vidlizer doctor          # shows ffmpeg, .env, provider, mlx-whisper status
vidlizer doctor --fix    # interactive repair: brew-installs ffmpeg/Ollama/LM Studio, re-runs setup

Manual provider setup

Ollama (fully offline, no API key):

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.