Import from YouTube and other sites via yt-dlp - #30
Merged
Merged
Conversation
Paste a link under the drop zone and a background job downloads it, catalogues it through the upload pipeline, and fills in what the site said: title, description, channel, date, licence, tags, chapter clips, and — with no Deepgram key — the captions as the transcript. Playlists fan out into one job per video. Closes the "URL import (M1) was specified and never built" gap in plan-of-attack. - Any site yt-dlp names, with its generic extractor excluded; the URL is checked against safe_url by the router and again by the job, since playlist entries are URLs the site chose. - Uploader tags applied by default (per-import checkbox routes them to suggestions); chapters become non-destructive clips; audio-only option; 1080p cap preferring H.264, configurable. - Attribution mapped from the site and stamped `embedded`; creator is never copied from the channel; retrieved_at finally has a writer. - ingest_upload split so a sync ingest_file shares the same _catalogue, and store_transcript split out so captions and Deepgram write alike. - yt-dlp[default,deno] pinned: YouTube now needs a JS runtime. Specified with decisions, operating notes and known limitations in docs/url-import.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017KYsSofkyeywy1NNPV8NtM
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Paste a link under the drop zone and a background job downloads it with yt-dlp. It then catalogues the file through the same pipeline an upload takes and fills in what the site said about it. This closes the "URL import (M1) was specified and never built" entry in
plan-of-attack.md. The full spec, decisions, operating notes and known limitations are indocs/url-import.md.What it does
genericextractor, which is the one that scrapes arbitrary pages.embeddedinfield_provenance:publisher← channelpublished_date,license,source_urlsource_title← series / album / titleretrieved_at, which finally has a writercreatoris never copied from the channel, following M10's "a blank beats a guess"URL_IMPORT_MAX_PLAYLIST_ITEMS. Videos already imported are skipped..m4a); a 1080p cap preferring H.264 (URL_IMPORT_MAX_HEIGHT); an optionalURL_IMPORT_COOKIES_FILE; duplicates finish pointing at the existing asset.How it's built
ingest/ytdlp.py: the only module that importsyt_dlpingest/captions.py: pure VTT parser that undoes YouTube's rolling duplicate linesenrichment/import_url.py: the jobrouters/imports.py:POST /api/assets/importservices/assets.py:ingest_uploadwas split so a syncingest_fileshares the same_catalogue. An import is probed, thumbnailed, indexed and queued for transcription by the upload's own code.enrichment/transcribe.py:store_transcriptwas split out so captions and Deepgram write transcripts identically.KIND_IMPORT_URLis a library kind, since no asset exists at submit time. The request rides inpayload, andcreated_asset_idis merged in on completion, soresult_asset_idworks as it does for M7 extractions. No migration.safe_urlin the router and again in the job, because playlist entries are URLs the site chose. Downloads use fixed filenames.UrlImport.tsxsits underUploadZone.LibraryViewadds a finished import and its chapter clips to the grid when it sees the job go from running to done, so old imports aren't re-surfaced on mount. The activity label is "Importing".Things to know
yt-dlp[default,deno]installs it from PyPI (manylinux x86_64 and aarch64, about 40 MB of image). The backend logs a warning at startup if it's missing.backend/requirements.txtand rebuild. The error message says so.URL_IMPORT_COOKIES_FILE. yt-dlp gets a copy of the file, because it writes the jar back out.Testing
pytest -q→ 1026 passed (949 existing + 77 new:test_url_import.py,test_imports_api.py,test_captions.py). yt-dlp is replaced by a fake that writes real fixture media, so ingest, ffprobe, thumbnails, clipping, the transcript writer and FTS all run for real. The options passed to yt-dlp are asserted on.YoutubeDLconstructor: postprocessors land in the right stages, and it detects Deno 2.9.7.format:check,lint,npm test(293 passed: 283 existing + 10 new), andnpm run buildare all clean.docker compose up --build -d.🤖 Generated with Claude Code
https://claude.ai/code/session_017KYsSofkyeywy1NNPV8NtM
Generated by Claude Code