Skip to content

Import from YouTube and other sites via yt-dlp - #30

Merged
davior merged 1 commit into
mainfrom
claude/brave-keller-1soxwa
Sep 28, 2026
Merged

davior merged 1 commit into
mainfrom
claude/brave-keller-1soxwa

Conversation

@davior

@davior davior commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Paste a link under the drop zone and a background job downloads it with yt-dlp. It then catalogues the file through the same pipeline an upload takes and fills in what the site said about it. This closes the "URL import (M1) was specified and never built" entry in plan-of-attack.md. The full spec, decisions, operating notes and known limitations are in docs/url-import.md.

What it does

  • Sites: any site yt-dlp names (YouTube, Rumble, Odysee, BitChute, Vimeo, X, podcasts…), except its generic extractor, which is the one that scrapes arbitrary pages.
  • Metadata: fields are mapped from the site and stamped embedded in field_provenance:
    • title, description
    • publisher ← channel
    • published_date, license, source_url
    • source_title ← series / album / title
    • retrieved_at, which finally has a writer
    • creator is never copied from the channel, following M10's "a blank beats a guess"
  • Tags: the uploader's tags are applied by default. A per-import checkbox sends them to the suggestion queue instead.
  • Chapters → clips: one non-destructive clip per chapter, clamped to the file's real length (per-import checkbox).
  • Captions as the transcript, only when no Deepgram key is set. Hand-written captions are preferred over automatic ones, and the original language over YouTube's auto-translations. With a key, the existing upload chain queues Deepgram as usual. Captions are keyword-indexed and chain an embed job.
  • Playlists fan out into one job per video, capped at URL_IMPORT_MAX_PLAYLIST_ITEMS. Videos already imported are skipped.
  • Other options: audio-only (.m4a); a 1080p cap preferring H.264 (URL_IMPORT_MAX_HEIGHT); an optional URL_IMPORT_COOKIES_FILE; duplicates finish pointing at the existing asset.

How it's built

  • Backend modules:
    • ingest/ytdlp.py: the only module that imports yt_dlp
    • ingest/captions.py: pure VTT parser that undoes YouTube's rolling duplicate lines
    • enrichment/import_url.py: the job
    • routers/imports.py: POST /api/assets/import
  • Refactors:
    • services/assets.py: ingest_upload was split so a sync ingest_file shares the same _catalogue. An import is probed, thumbnailed, indexed and queued for transcription by the upload's own code.
    • enrichment/transcribe.py: store_transcript was split out so captions and Deepgram write transcripts identically.
  • Job shape: KIND_IMPORT_URL is a library kind, since no asset exists at submit time. The request rides in payload, and created_asset_id is merged in on completion, so result_asset_id works as it does for M7 extractions. No migration.
  • Security: the URL is checked by safe_url in the router and again in the job, because playlist entries are URLs the site chose. Downloads use fixed filenames.
  • Frontend: UrlImport.tsx sits under UploadZone. LibraryView adds a finished import and its chapter clips to the grid when it sees the job go from running to done, so old imports aren't re-surfaced on mount. The activity label is "Importing".

Things to know

  1. yt-dlp needs Deno for YouTube now. yt-dlp[default,deno] installs it from PyPI (manylinux x86_64 and aarch64, about 40 MB of image). The backend logs a warning at startup if it's missing.
  2. YouTube breaks yt-dlp every few weeks. When imports start failing, bump the pin in backend/requirements.txt and rebuild. The error message says so.
  3. Bot checks on datacenter IPs. YouTube often shows "Sign in to confirm you're not a bot" there; use URL_IMPORT_COOKIES_FILE. yt-dlp gets a copy of the file, because it writes the jar back out.
  4. A video with chapter clips can't be deleted directly. The M7 delete guard blocks it until the clips are removed or promoted.
  5. Long playlists fill the grid on the next load, not live. The activity feed only shows 10 rows.

Testing

  • Backend: pytest -q → 1026 passed (949 existing + 77 new: test_url_import.py, test_imports_api.py, test_captions.py). yt-dlp is replaced by a fake that writes real fixture media, so ingest, ffprobe, thumbnails, clipping, the transcript writer and FTS all run for real. The options passed to yt-dlp are asserted on.
  • The built options were also checked against the real YoutubeDL constructor: postprocessors land in the right stages, and it detects Deno 2.9.7.
  • Frontend: format:check, lint, npm test (293 passed: 283 existing + 10 new), and npm run build are all clean.
  • Not tested against real sites. This sandbox's proxy blocks youtube.com. The first real import (one video with chapters, one audio-only, one playlist, one Rumble/Odysee link) needs to be done by hand after docker compose up --build -d.

🤖 Generated with Claude Code

https://claude.ai/code/session_017KYsSofkyeywy1NNPV8NtM


Generated by Claude Code

Paste a link under the drop zone and a background job downloads it,
catalogues it through the upload pipeline, and fills in what the site
said: title, description, channel, date, licence, tags, chapter clips,
and — with no Deepgram key — the captions as the transcript. Playlists
fan out into one job per video. Closes the "URL import (M1) was
specified and never built" gap in plan-of-attack.

- Any site yt-dlp names, with its generic extractor excluded; the URL is
  checked against safe_url by the router and again by the job, since
  playlist entries are URLs the site chose.
- Uploader tags applied by default (per-import checkbox routes them to
  suggestions); chapters become non-destructive clips; audio-only option;
  1080p cap preferring H.264, configurable.
- Attribution mapped from the site and stamped `embedded`; creator is
  never copied from the channel; retrieved_at finally has a writer.
- ingest_upload split so a sync ingest_file shares the same _catalogue,
  and store_transcript split out so captions and Deepgram write alike.
- yt-dlp[default,deno] pinned: YouTube now needs a JS runtime.

Specified with decisions, operating notes and known limitations in
docs/url-import.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017KYsSofkyeywy1NNPV8NtM
@davior
davior marked this pull request as ready for review September 28, 2026 09:28
@davior
davior merged commit c1c2878 into main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants