Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P1w53ybPWArpZpcei9JPYX
87 lines
2.5 KiB
Markdown
87 lines
2.5 KiB
Markdown
# yt-dlf
|
|
|
|
YouTube Downloader & Transcript Extractor — a local web frontend for `yt-dlp` and `ffmpeg`.
|
|
|
|
## Features
|
|
|
|
- Paste a YouTube URL and download the best available video quality
|
|
- Subtitles downloaded automatically in the video's original language, English, and German (where available)
|
|
- Subtitles converted to clean Markdown (timestamps stripped, text deduplicated and paragraph-wrapped)
|
|
- Videos without subtitles (Instagram, archive.org, …) are transcribed locally via Whisper (faster-whisper, GPU-accelerated) — see `INSTALL.md`
|
|
- Optional audio extraction to MP3 via ffmpeg
|
|
- Real-time progress log streamed to the browser
|
|
- Files saved to `~/YouTube/<video title>/`
|
|
|
|
## Prerequisites
|
|
|
|
```bash
|
|
brew install yt-dlp ffmpeg
|
|
```
|
|
|
|
Node.js 16+ is required for the web app (tested with Node 26).
|
|
|
|
## Project Layout
|
|
|
|
```
|
|
yt-dlf/
|
|
├── scripts/
|
|
│ └── subtitle_to_markdown.py # standalone CLI converter
|
|
├── src/
|
|
│ ├── lib/
|
|
│ │ └── subtitle.js # JS conversion module (used by web app)
|
|
│ └── routes/
|
|
│ ├── api/download/
|
|
│ │ └── +server.js # SSE endpoint — runs yt-dlp / ffmpeg
|
|
│ ├── +layout.svelte
|
|
│ └── +page.svelte # UI
|
|
└── vite.config.js
|
|
```
|
|
|
|
## Running the web app
|
|
|
|
```bash
|
|
npm install # first time only
|
|
npm run dev # starts on http://localhost:5173
|
|
```
|
|
|
|
To use a custom port:
|
|
|
|
```bash
|
|
PORT=8080 npm run dev
|
|
```
|
|
|
|
## Standalone transcription
|
|
|
|
Transcribe any local video/audio file to WebVTT directly from the terminal
|
|
(requires the faster-whisper venv, see `INSTALL.md`):
|
|
|
|
```bash
|
|
.venv/bin/python scripts/transcribe.py video.mp4
|
|
```
|
|
|
|
The language is detected automatically and the subtitle file is written next
|
|
to the input as `video.<lang>.vtt` (e.g. `video.de.vtt`). Model and device can
|
|
be overridden via environment variables:
|
|
|
|
```bash
|
|
WHISPER_MODEL=large-v3 .venv/bin/python scripts/transcribe.py video.mp4 # best quality
|
|
WHISPER_DEVICE=cpu .venv/bin/python scripts/transcribe.py video.mp4 # don't touch the GPU
|
|
```
|
|
|
|
To get a Markdown transcript, feed the VTT through the subtitle converter below:
|
|
|
|
```bash
|
|
./scripts/subtitle_to_markdown.py video.de.vtt
|
|
```
|
|
|
|
## Standalone subtitle converter
|
|
|
|
Convert a `.vtt` or `.srt` subtitle file to Markdown directly from the terminal:
|
|
|
|
```bash
|
|
./scripts/subtitle_to_markdown.py video.en.vtt
|
|
./scripts/subtitle_to_markdown.py video.en.vtt output.md
|
|
```
|
|
|
|
Output is saved next to the input file (same name, `.md` extension) unless a second argument is given.
|