# Social video downloader One command for Instagram and TikTok, with batch files, proxy routing, resumable media, and a searchable directory of **200 entries (186 unique platform/account pairs)**. ## Quick start Unzip the toolkit and open a terminal in its folder. Python 3.10+ is required. On macOS/Linux: ```sh chmod +x download setup.sh ./download instagram:leoniehanne tiktok:leoniehanne --plan ./download instagram:leoniehanne tiktok:leoniehanne ``` The launcher installs dependencies in a local `.venv` on first use. A plan makes **no social-platform requests**; first-time dependency installation still accesses the package index. You can run `./setup.sh` separately. For Windows or a manual installation: ```sh python -m venv .venv # Windows: .venv\Scripts\python -m pip install -r requirements.txt .venv\Scripts\python social_download.py tiktok:leoniehanne --plan # macOS/Linux: .venv/bin/python social_download.py tiktok:leoniehanne --plan ``` By default each account selects **30 newest videos + 10 popular videos from the last three calendar months**, downloading overlaps once. Instagram includes carousel video clips; each clip counts as one video. Photo-only posts, TikTok photo slideshows, Stories, and live streams are outside this command's scope. The existing Instagram HAR archive tool remains available for mixed photo/video post downloads; see README.md. ## Multiple accounts Open `influencer-shortlist.html`, filter by platform/category or creator name, and click **Export account JSON**. Use the downloaded file: ```sh ./download --accounts-file ~/Downloads/accounts.json --plan ./download --accounts-file ~/Downloads/accounts.json --resume ``` Or edit `accounts.example.json`. Duplicate accounts are removed before processing. Each account runs sequentially, failures are recorded, and the batch proceeds to the next account. `--fail-fast` stops after the first failure or incomplete result. Use the bundled directory directly: ```sh # Preview the 25 TikTok fashion accounts ./download --directory --category fashion --platform tiktok --plan # Start with one account ./download --directory --category fashion --platform tiktok --limit-accounts 1 # Process the whole filtered list ./download --directory --category fashion --platform tiktok --resume ``` Running `--directory` without filters selects all **186 unique platform/account pairs**. It does not start automatically when opening the HTML or installing the toolkit. Category membership is editorial, not a measured global ranking. `influencers.json` is the canonical directory; run `python3 build_directory.py` after editing it. ## Selection, limits, and request volume ```sh ./download tiktok:wisdm8 --recent 30 --popular 10 --since 2026-07-02 ./download instagram:leoniehanne --max-scan 500 --max-requests 300 ./download tiktok:leoniehanne --dry-run ``` `--plan` is offline. `--dry-run` performs metadata requests and saves selections without transferring media. Latest selection uses publish timestamps, not profile grid order. The popularity cutoff only applies to the popular list; recent videos may be older than the cutoff. Default scan cap: **200 posts per source**. Instagram has separate feed and Reels sources (up to 400 listing entries before deduplication); TikTok has one profile source. A single look-ahead entry may be read to detect truncation. The tool does not stop on an old pinned post. A capped scan is explicitly reported as **partial**; it does not claim a complete three-month ranking merely because an old post appeared. Default request cap: **150 HTTP attempts per account**, including metadata and media. Requests are spaced by at least **1 second**, with **5 seconds between accounts**. These are adjustable using `--max-requests`, `--request-delay`, and `--account-delay`. Underlying platform rate limiting can wait longer. yt-dlp counts each `urlopen` invocation; redirects handled inside its networking backend can add wire-level requests. Instaloader requests count each Requests `send`, including redirects. This is an application request budget, not an exact packet count or time limit. For 30 recent + 10 popular, there are at most **40 selected video files per account**. TikTok typically needs profile/pagination requests plus roughly one detail lookup and one transfer per selected video; actual pagination, retries, redirects and request counts vary. Instagram may need additional detail lookups to expose metrics. The hard budget stops requests rather than guessing that a profile will fit an estimate. `--plan` prints the aggregate configured budget before you run a batch. Auto ranking chooses one metric with the best coverage in the popularity window, preferring **views → plays → likes + comments → likes → comments**, then shares and saves if those are the only available metrics. TikTok exposes shares/saves when available. You can force `--rank-by shares` or another metric. Instagram's adapter does not expose shares/saves. Hidden counts stay unknown; unavailable scores are excluded and reported. The tool does not combine views and likes into one uncalibrated score. ## Access and proxies Public, unsigned access is the default. No automatic browser-cookie reading or sign-in takes place. Profile URLs and handles work on both platforms; numeric Instagram IDs also work. TikTok numeric user IDs alone are not supported by this adapter: supply the public username. ```sh export SOCIAL_PROXY='http://HOST:PORT' ./download --accounts-file accounts.example.json ``` An explicit `--proxy` overrides `SOCIAL_PROXY`, then `INSTAGRAM_PROXY`. HTTP(S), SOCKS5 and SOCKS5h are supported; credentials should be percent-encoded and kept in environment variables rather than shared commands. No proxy provider is bundled. A configured proxy applies to both metadata and media and does not fail open to a direct connection. A proxy does not guarantee access; the tool does not rotate identities after a block. Optional existing authorized sessions: ```sh ./download instagram:leoniehanne --login YOUR_LOGIN --session-file .sessions/session-YOUR_LOGIN ./download tiktok:leoniehanne --cookies .sessions/tiktok-cookies.txt ``` TikTok accepts a **Netscape-format cookie file**, not a browser database. Supplied cookies are used for metadata extraction; media then uses a separate unsigned client. With no cookie file, TikTok retains the public extractor's generated session cookies for CDN compatibility; this session is not signed in. If a CDN requires login, the transfer fails rather than silently forwarding login cookies. Instagram likewise uses an anonymous session for media bytes. Keep session and cookie files local and private. Only load trusted Instaloader session files (they use pickle). ## Files and recovery ```text downloads/ instagram/USERNAME/manifest.json instagram/USERNAME/videos/*.mp4 tiktok/USERNAME/manifest.json tiktok/USERNAME/videos/*.mp4 batches/BATCH_ID.json ``` Each manifest contains ordered `recent`/`popular` video IDs, metric counts, coverage, status, paths, byte sizes and SHA-256 hashes. Files are published atomically after transfer; signed CDN URLs, cookies and proxy credentials are excluded from manifests. Earlier selections stay on disk. Reruns verify completed media and skip intact files. `--resume` additionally skips accounts marked complete in the matching batch checkpoint after checking that their selected media is still intact. Partial/failed accounts and accounts with missing/corrupt files are retried. Use the same explicit `--since YYYY-MM-DD` to resume a long batch across different dates; the default cutoff uses UTC dates and changes daily. Omit `--resume` when you want to refresh metrics and newest selections for every account. The tool stops downloads for an account after the first network/extraction failure, preserving earlier files. The next run retries unfinished files. Do not run concurrent jobs against the same output: per-account write locks protect manifests. If an ungraceful termination leaves `.download.lock`, remove it only after confirming that no download is active. Exit codes: **0** successful/plan, **1** invalid setup, **2** failed or partial batch, **130** interrupted. A partial result may mean the scan cap was reached, some metrics were missing, too few videos existed, or a transfer failed. Inspect the manifest warnings and batch statuses. ## Troubleshooting If public access is rejected, complete any platform verification normally and retry later, or explicitly supply an authorized session. Anonymous profile enumeration is not guaranteed on either platform. Failed metadata scans never become successful empty selections. An extraction error can also mean an unsupported photo slideshow or an upstream endpoint change. The toolkit prints exception types and safe hints instead of raw upstream diagnostics that can contain signed URLs or credentials. Dependencies are pinned where endpoint compatibility matters. To upgrade the TikTok extractor deliberately: change the yt-dlp pin in `requirements.txt` to a tested release, then rerun `./setup.sh`. `curl-cffi` provides the browser-impersonation transport used by current yt-dlp extractors; it does not import a browser session. The tool prefers combined H.264 MP4 for compatibility, falling back to another combined MP4 format. FFmpeg is unnecessary for these downloads. Live validation: anonymous TikTok profile metadata and one Leonie Hanne MP4 transfer succeeded during development. This verifies that path at the time of the test, not every account or future availability. Explicit-cookie isolation and both proxy transports are covered by local tests. No live proxy-provider test or bulk directory download was performed. Run offline tests with `.venv/bin/python -m unittest discover -s tests -v`. Rebuild the portable package with `python3 build_package.py`; it uses an allowlist and excludes media, cookies, sessions, and virtual environments. Implementation references: [yt-dlp](https://github.com/yt-dlp/yt-dlp), [TikTok extractor](https://github.com/yt-dlp/yt-dlp/blob/master/yt_dlp/extractor/tiktok.py), [Instaloader API](https://instaloader.github.io/module/structures.html).