Skip to writing
Soheil
All writing
aiautomationmacollamabashtechnical

Building a Local YouTube Summarizer with Mac Shortcuts, Ollama, and Qwen: Zero Subscriptions, Full Privacy

How I built a completely offline YouTube video summarizer using Mac Shortcuts, Ollama, yt-dlp, and the open-source Qwen model—no cloud, no subscriptions, just local AI doing the heavy lifting.

Ever find yourself watching hour-long YouTube talks just to get the main points? I did too. I kept thinking there must be a better way to get the summary without sitting through the whole thing. Then I came across a paid service that does exactly that, but instead of subscribing, I decided to build it myself.

I put together a quick setup using Mac Shortcuts, Ollama (running the open-source Qwen model), and a small bash script. The Shortcut takes a YouTube link, grabs the captions with yt-dlp, cleans them up, and then uses the local Qwen model to produce a structured Markdown summary with sections like TL;DR, key points, and takeaways.

It runs entirely locally: no cloud, no subscriptions, and it works surprisingly well.


The Shortcut (at a glance)

Mac Shortcuts workflow showing the 5 actions


How It Works

  1. Captions – yt-dlp downloads English (then Farsi) captions where available.
  2. Cleaning – A tiny script strips timestamps/tags and de-duplicates lines.
  3. Summarization – The cleaned transcript is sent to Qwen (via Ollama) locally.
  4. Output – A .summary.md file lands in: ~/Library/Application Support/ytsum/

Install the Tools (once)


/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"

# Utilities
brew install yt-dlp jq

# Local AI runtime
brew install --cask ollama
open -a Ollama               # starts the daemon

# Model (pick one; 14B is best if your Mac can handle it)
ollama pull qwen3:14b
# or: ollama pull qwen3:8b
# or: ollama pull qwen3:4b

Add the Script

Create ~/bin/yt_summary.sh, make it executable, and you’re done.

mkdir -p ~/bin
nano ~/bin/yt_summary.sh

Update the script content (feel free to adjust the prompt):

#!/usr/bin/env bash
# YouTube summarizer using captions + Ollama/Qwen3
# Works offline, respects privacy, zero subscriptions

set -euo pipefail
set -f

OUTDIR="${OUTDIR:-$HOME/Library/Application Support/ytsum}"
mkdir -p "$OUTDIR"
LOGFILE="$OUTDIR/yt_summary_$(date +%Y%m%d-%H%M%S).log"
DEBUG="${DEBUG:-1}"

# Logging (stderr -> LOGFILE; only JSON to stdout)
exec 2>>"$LOGFILE"
[[ "$DEBUG" == "1" ]] && set -x
log() { printf '%s\n' "$*" >>"$LOGFILE"; }
die() { log "FATAL: $*"; printf '{"SUMMARY_MD":"%s","SUMMARY_MD_URL":"file://%s"}\n' "$LOGFILE" "$LOGFILE"; exit 0; }
trap 'die "Error at line $LINENO (see log)."' ERR

# PATH
BREW_PREFIX="$(/usr/bin/env brew --prefix 2>/dev/null || true)"
[[ -z "${BREW_PREFIX:-}" ]] && { [[ -d /opt/homebrew ]] && BREW_PREFIX=/opt/homebrew || BREW_PREFIX=/usr/local; }
export PATH="$BREW_PREFIX/bin:/usr/local/bin:/opt/homebrew/bin:/usr/bin:/bin:/usr/sbin:/sbin"

command -v yt-dlp >/dev/null || die "yt-dlp not found"
command -v ollama >/dev/null || die "ollama not found"
command -v jq >/dev/null || die "jq not found"

URL="${1:-}"; [[ -z "$URL" ]] && die "No URL passed"

# VTT → clean text
vtt_to_txt() {
  /usr/bin/sed -E '/^WEBVTT|^Kind:|^Language:|^[[:space:]]*$|-->/d' "$1"   | /usr/bin/sed -E 's#</?c>##g; s#<[0-9:.]+>##g; s#<[^>]+>##g'   | /usr/bin/awk '{ gsub(/^[ \t]+|[ \t]+$/, "", $0); if ($0 != "" && $0 != prev) print; prev=$0 }'
}

# Ensure Qwen3 model (tries 14B, 8B, 4B)
ensure_qwen3() {
  if ! pgrep -x "ollama" >/dev/null; then
    log "Starting Ollama…"; open -gj -a "Ollama" || true; sleep 3
  fi
  local order=("qwen3:14b" "qwen3:8b" "qwen3:4b")
  local model=""
  for m in "${order[@]}"; do
    if ollama list 2>/dev/null | grep -qE "^[[:space:]]*$m([[:space:]]|$)"; then model="$m"; break; fi
  done
  if [[ -z "$model" ]]; then
    model="qwen3:14b"
    log "Pulling $model…"; ollama pull "$model" || { model="qwen3:8b"; ollama pull "$model" || { model="qwen3:4b"; ollama pull "$model"; }; }
  fi
  echo "$model"
}

# Prompt: monologue → first person; debate → third person w/ attribution
make_prompt() {
cat <<'PROMPT'
Use only the transcript. Do NOT repeat any instruction text. Output ONLY between <<<BEGIN_MD>>> and <<<END_MD>>>.

VOICE RULE:
- If the transcript is a single-speaker monologue (no clear turn-taking, no speaker labels like "HOST:", "GUEST:", ">> Name"), write summaries in FIRST PERSON as the speaker.
- If it's a discussion/debate with multiple speakers, write in NEUTRAL THIRD PERSON and attribute statements (e.g., "Host: …", "Guest: …") when clear.

QUALITY RULES:
- Preserve proper nouns, dates, figures, and claims exactly as stated.
- Be faithful to the source—no external facts, no speculation.
- De-duplicate repeated lines common in auto-captions.
- Prefer chronological flow for the long summary; explain context and conclusions.

<<<BEGIN_MD>>>
## Short TL;DR
- 6–8 concise bullets for core ideas & outcomes

## Long Summary
- 350–500 words; high-recall narrative of the full arc (setup → key arguments/evidence → conclusions/implications)
- If monologue: write in first person as the speaker. If multi-speaker: third person with who-said-what when identifiable.

## Participants
- Named people/groups with stated roles (if any)
- If none named: "Speaker not named"

## Key Points & Numbers
- 5–10 bullets of notable facts/claims/figures/dates from the transcript

## Open Questions / Next Steps
- 2–5 bullets for unresolved items or proposed actions
<<<END_MD>>>
PROMPT
}

run_qwen3() {
  local model="$1"
  { make_prompt; printf "\nTRANSCRIPT:\n%s\n" "$TRANSCRIPT"; } | ollama run "$model"
}

# Resolve video/meta
VID="$(yt-dlp --ignore-config --cookies-from-browser chrome --get-id "$URL" 2>>"$LOGFILE" || true)"
[[ -z "${VID:-}" ]] && die "Could not resolve video id"
TITLE="$(yt-dlp --ignore-config --cookies-from-browser chrome --get-title "$URL" 2>>"$LOGFILE" || true)"
[[ -z "${TITLE:-}" ]] && TITLE="$VID"
BASENAME="$OUTDIR/$VID"

# Captions: auto EN → auto FA → official EN
TRANSCRIPT=""
log "Captions: auto EN"
yt-dlp --ignore-config --cookies-from-browser chrome   --write-auto-sub --skip-download --sub-lang en   --sub-format vtt --convert-subs vtt -o "$OUTDIR/%(id)s.%(ext)s" "$URL" >>"$LOGFILE" 2>&1 || true
[[ -f "$OUTDIR/$VID.en.vtt" ]] && TRANSCRIPT="$(vtt_to_txt "$OUTDIR/$VID.en.vtt")"

if [[ -z "$TRANSCRIPT" ]]; then
  log "Captions: auto FA"
  yt-dlp --ignore-config --cookies-from-browser chrome     --write-auto-sub --skip-download --sub-lang fa     --sub-format vtt --convert-subs vtt -o "$OUTDIR/%(id)s.%(ext)s" "$URL" >>"$LOGFILE" 2>&1 || true
  [[ -f "$OUTDIR/$VID.fa.vtt" ]] && TRANSCRIPT="$(vtt_to_txt "$OUTDIR/$VID.fa.vtt")"
fi

if [[ -z "$TRANSCRIPT" ]]; then
  log "Captions: official EN"
  yt-dlp --ignore-config --cookies-from-browser chrome     --write-sub --skip-download --sub-lang en     --sub-format vtt --convert-subs vtt -o "$OUTDIR/%(id)s.%(ext)s" "$URL" >>"$LOGFILE" 2>&1 || true
  [[ -f "$OUTDIR/$VID.en.vtt" ]] && TRANSCRIPT="$(vtt_to_txt "$OUTDIR/$VID.en.vtt")"
fi

[[ -z "$TRANSCRIPT" ]] && die "No captions available (en/fa)."

# Truncate long transcripts (keeps model fast)
MAX_CHARS=160000
if (( ${#TRANSCRIPT} > MAX_CHARS )); then
  TRANSCRIPT="${TRANSCRIPT:0:MAX_CHARS}"$'\n\n[... truncated for length ...]'
fi

# Summarize
MODEL="$(ensure_qwen3)"
RAW="$(run_qwen3 "$MODEL")"

# Extract the marked block; fallback to raw if markers missing
MD_BLOCK="$(printf "%s\n" "$RAW" | sed -n '/^<<<BEGIN_MD>>>$/,/^<<<END_MD>>>$/p' | sed '1d;$d')"
[[ -z "$MD_BLOCK" ]] && MD_BLOCK="$RAW"

# Save
MD_PATH="${BASENAME}.summary.md"
printf "# %s\n\n%s\n" "$TITLE" "$MD_BLOCK" > "$MD_PATH"

# Return JSON for Shortcuts
printf '{"SUMMARY_MD":"%s","SUMMARY_MD_URL":"file://%s","MODEL":"%s"}\n' "$MD_PATH" "$MD_PATH" "$MODEL"

Save and exit:

chmod +x ~/bin/yt_summary.sh

What's Next

I'm considering a few improvements:

  • Adding Whisper as a fallback for videos without captions (so it can transcribe audio locally)
  • Batch mode for playlists—paste a playlist URL, get summaries for every video
  • A simple desktop UI (maybe Electron or Tauri) for people who don't want to deal with Shortcuts
  • Smart caching so repeated videos just load the existing summary instead of re-processing

For now though, this does exactly what I needed: one click, one summary, zero subscriptions. If you end up building something similar or have ideas for improvements, send me a note. Always curious to see what people do with this kind of setup.

Back to the notebook