跳到資源內容
回到全部資源
閱讀+複製可直接閱讀保留 Markdown 下載

給 AI 助理的 Colab 說明書

把 Google AI Pro 每月 200 點 Colab 額度交給你的 AI 助理跑:安裝登入、安全跑法(背景執行、不斷線空轉扣點)、成本實測與踩坑排除,整份貼給它就會跑。

發布:

下載 .md

可以直接往下閱讀;遇到指令或提示詞,右上角可以單獨複製。

有 Google AI Pro 的話,你每個月有 200 點 Colab 額度,可以借 Google 機房裡的 GPU 電腦跑 AI 模型:生影片、轉大量逐字稿、訓練自己的小模型都行。

這份說明書是寫給你的 AI 助理(Claude Code、Codex 這類能在你電腦上執行指令的助理)看的。整份貼給它,它照著做就能開機、跑工作、拿回結果、關機,不用自己從頭研究,也就不會在研究這一步花掉你的用量。

我自己是讓 Claude Code 照這套流程,用不到 1 點生出一支 8 秒、會講中文的卡通短片。

你要自己做的只有兩件事

  1. 把這整份說明書貼給你的 AI 助理,告訴它你想跑什麼(例如「幫我用 Colab 生一支 8 秒的影片」)。
  2. 第一次登入時,它會在終端機給你一串 Google 網址。你用瀏覽器打開、登入有 AI Pro 的那個帳號、按同意,再把畫面上的驗證碼貼回終端機。

其他都交給它。

先知道的規則

  • 只有機器開著才扣點。A100 高記憶體實測每小時 6.77 點,200 點大約能開 29 個小時。關機就停止計費。
  • 這份額度只有方案管理者本人能用,家庭成員分不到;要滿 18 歲。
  • 點數每個月發一次(Colab 官方說明:「a monthly allotment of Colab compute units」)。我在 Colab 的方案頁上看到的說明是點數 90 天後失效。
  • 實測:已經有一台 GPU 機器在跑時,再開第二台會被拒絕(CPU 機器可以另外開)。
  • Colab 條款禁止拿它架長期對外的網站或 API。只能「開機、跑完、關機」。

以下給 AI 助理

1. 安裝與登入

bash
# 安裝 Google 官方的 Colab 指令列工具(Linux 與 macOS)
uv tool install google-colab-cli
colab version

# 第一次登入:用 oauth2,不要用 ADC(重做 ADC 會蓋掉使用者既有的 gcloud 憑證)
# 這一步要請人類在瀏覽器完成授權、把驗證碼貼回終端機
colab --auth=oauth2 usage

uv tool install 如果很久沒動,多半是別的 uv 程序佔著快取鎖。換一個獨立的快取資料夾再裝:UV_CACHE_DIR=/tmp/uv-colab uv tool install google-colab-cli。

登入成功會看到 Current balance: 200.00 compute units。之後每個 colab 指令都要帶 --auth=oauth2。

2. 安全跑法:不要用一條連線跑完長工作

最重要的一條。 用一次 colab exec 跑一個好幾分鐘的工作,本機網路一抖,連線就會斷(RuntimeError: Connection was lost.),而且指令會卡住不退出,雲端機器卻照樣計費。我第一次就這樣空轉了 10 分鐘。

正確做法:

  1. 開機後先做一次暖身 exec,再上傳檔案(開機後直接上傳、才第一次 exec,會斷線)。
  2. 在雲端用 subprocess.Popen(..., start_new_session=True) 把工作丟到背景跑,輸出寫進 log。
  3. 本機每 45 秒用很短的 exec 查一次 log。查的那一次斷線沒關係,下一次再查。
  4. 每個 colab 指令都設逾時,卡住就砍掉重試。
  5. 用 try/finally 保證最後一定 colab stop,另外設總時長上限。
  6. 開機前先 colab sessions 檢查同名機器,不要重複開。

下面這支工具已經把上面六條都做好了。存成 colab_detached.py,把要在雲端跑的 Python 腳本當參數丟給它:

bash
# 例:在 A100 上跑 job.py,先上傳 input.jpg,跑完下載 /content/out.mp4
python3 colab_detached.py job.py --gpu A100 --high-mem --upload input.jpg --download /content/out.mp4 --max-minutes 60
python
"""Run one long job on a Google Colab GPU without babysitting a fragile connection.

Why this exists: a single long `colab exec` keeps one websocket open for the whole job.
If your home network blips, the CLI can lose the connection and hang forever while the
VM keeps billing. This script instead starts the job *detached* on the VM, checks its
log with short commands, downloads the results, and always stops the VM.

Usage:
  python3 colab_detached.py job.py --gpu A100 --upload input.jpg --download /content/out.mp4
  python3 colab_detached.py job.py --gpu T4 --max-minutes 30

Requirements: the official Colab CLI (`uv tool install google-colab-cli`), already logged in
with `colab --auth=oauth2 usage`. Tested with Colab CLI 0.7.4 on macOS, 2026-09-28.
"""
import argparse
import subprocess
import sys
import time
from pathlib import Path

AUTH = "--auth=oauth2"


def colab(args, stdin=None, timeout=120):
    """Run one colab CLI command with a hard timeout, so a dropped connection can't hang us."""
    try:
        r = subprocess.run(["colab", AUTH, *args], input=stdin, capture_output=True, text=True, timeout=timeout)
        return r.returncode, (r.stdout + r.stderr).strip()
    except subprocess.TimeoutExpired:
        return -9, f"timed out after {timeout}s"


def remote(session, code, timeout=90):
    return colab(["exec", "-s", session], stdin=code, timeout=timeout)


def usage_line():
    _, out = colab(["usage"], timeout=60)
    return " | ".join(out.splitlines()[-3:])


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("job", help="local Python script to run on the VM (runs from /content)")
    ap.add_argument("--gpu", default="T4", help="CPU, T4, L4, G4, A100 or H100 (A100 High-RAM measured 6.77 units/hour)")
    ap.add_argument("--high-mem", action="store_true", help="request the high-RAM machine shape")
    ap.add_argument("--session", default="detached-job")
    ap.add_argument("--upload", action="append", default=[], help="local file to copy into /content (repeatable)")
    ap.add_argument("--download", action="append", default=[], help="remote file to fetch when the job ends (repeatable)")
    ap.add_argument("--max-minutes", type=float, default=60, help="hard wall-clock limit; the VM is stopped after this")
    a = ap.parse_args()

    # Only one GPU VM at a time is allowed on consumer plans; never reuse a live session by accident.
    _, sessions = colab(["sessions"], timeout=60)
    if any(line.startswith(f"[{a.session}]") for line in sessions.splitlines()):
        sys.exit(f"session {a.session} is already running; stop it first with: colab {AUTH} stop -s {a.session}")

    print("before:", usage_line(), flush=True)
    new_args = ["new", "-s", a.session] + ([] if a.gpu.upper() == "CPU" else ["--gpu", a.gpu]) + (["--high-mem"] if a.high_mem else [])
    rc, out = colab(new_args, timeout=600)
    print(out.splitlines()[-1] if out else out, flush=True)
    if rc != 0:
        sys.exit("could not allocate a VM (no charge). Try another --gpu or wait a few minutes.")

    t0 = time.monotonic()
    try:
        # Warm up the kernel BEFORE uploading: on 2026-09-28 the first exec after an upload
        # lost its connection twice in a row on A100 VMs; a warm-up exec first never did.
        for _ in range(4):
            rc, out = remote(a.session, "print('warm', 1 + 1)")
            if rc == 0 and "warm 2" in out:
                break
            time.sleep(20)
        else:
            sys.exit("kernel never answered")

        for f in [a.job, *a.upload]:
            rc, out = colab(["upload", "-s", a.session, f, f"/content/{Path(f).name}"], timeout=600)
            print(out.splitlines()[-1] if out else out, flush=True)
            if rc != 0:
                sys.exit(f"upload failed: {f}")

        job = f"/content/{Path(a.job).name}"
        start = (
            "import os, subprocess, sys\n"
            "if not os.path.exists('/content/job.pid'):\n"
            "    log = open('/content/job.log', 'a')\n"
            f"    p = subprocess.Popen([sys.executable, '-u', {job!r}], cwd='/content', stdout=log,\n"
            "                         stderr=subprocess.STDOUT, start_new_session=True)\n"
            "    open('/content/job.pid', 'w').write(str(p.pid))\n"
            "print('started', open('/content/job.pid').read())\n"
        )
        for _ in range(4):
            rc, out = remote(a.session, start)
            if rc == 0 and "started" in out:
                break
            time.sleep(20)
        else:
            sys.exit("could not start the job")
        print("job started; checking every 45 s", flush=True)

        poll = (
            "import os\n"
            "t = open('/content/job.log').read() if os.path.exists('/content/job.log') else ''\n"
            "pid = int(open('/content/job.pid').read())\n"
            "alive = os.path.exists(f'/proc/{pid}') and open(f'/proc/{pid}/stat').read().split()[2] != 'Z'\n"
            "tail = [l for l in t.splitlines() if l.strip()][-1:] or ['']\n"
            "print('ALIVE' if alive else 'EXITED', '|', tail[0][-160:])\n"
        )
        last = ""
        while True:
            if time.monotonic() - t0 > a.max_minutes * 60:
                print("wall-clock limit reached; stopping the VM", flush=True)
                return 1
            time.sleep(45)
            rc, out = remote(a.session, poll)
            status = out.splitlines()[-1] if out else ""
            if rc != 0:
                continue  # a dropped check is harmless; the job keeps running on the VM
            if status != last:
                print(f"[{(time.monotonic() - t0) / 60:.1f} min] {status}", flush=True)
                last = status
            if status.startswith("EXITED"):
                break

        _, tail = remote(a.session, "print(open('/content/job.log').read()[-3000:])")
        Path("job.log").write_text(tail)
        print("remote log saved to job.log", flush=True)
        for f in a.download:
            rc, out = colab(["download", "-s", a.session, f, Path(f).name], timeout=900)
            print(out.splitlines()[-1] if out else out, flush=True)
        return 0
    finally:
        rc, out = colab(["stop", "-s", a.session], timeout=300)
        print(out.splitlines()[-1] if out else out, flush=True)
        print(f"total {(time.monotonic() - t0) / 60:.1f} min;", "after:", usage_line(), flush=True)


if __name__ == "__main__":
    sys.exit(main())

3. 成本(2026-09-28 實測)

工作花費時間扣點
A100 高記憶體,每小時—6.77 點
影片模型第一次開機準備(裝軟體、下載約 55 GB 模型)約 4.5 分鐘約 0.5 點(每次開機付一次)
8 秒卡通短片(加速版 4 步),從開機到拿回影片8.6 分鐘約 1 點
10 秒直式短片,同一台機器連續生第二支約 5.6 分鐘約 0.6 點
11.5 秒短片,完整品質 20 步(含開機準備)約 33 分鐘約 3.7 點
CPU 機器,每小時—約 0.08 點

要省點,就一次開機、連續跑多個工作,把開機準備的時間攤掉。

4. 範例:生一支會講中文的短片(MiniMax H3)

MiniMax H3 是能生出「影片+同步聲音」的開放模型(官方模型卡:huggingface.co/MiniMaxAI/MiniMax-H3)。授權可以商用(年營收超過 2,000 萬美元要另外申請),但排除歐盟、英國、南韓、美國;公開發佈生成的內容時要清楚標示是 AI 生成,做成產品時畫面上要標示 MiniMax H3。

最省事的起點是 @killkli 在 Threads 分享的 Colab 技能(github.com/killkli/minimax-h3-colab-skill):裡面的 notebook 會在 Colab 裝好 ComfyUI、下載官方量化權重、跑「參考圖轉影片」。請 AI 助理讀那個 repo,再用第 2 節的背景跑法執行,不要用它預設的一條連線跑到底。

拍得好看的四個重點

  1. 首幀定錨:先用生圖工具畫好第一格,讓模型只負責「讓它動起來」。不給首幀,它會自己編構圖。
  2. 每個鏡頭只做一個動作,同一段畫風描述每一鏡逐字重貼。
  3. 先用加速版打草稿(4 步),確認動作方向對了,再用完整品質(20 步)出定稿。
  4. 解析度照官方上限:橫式 1344×768、直式 768×1344。原生就是 768p。

提示詞範本(照官方「參考圖轉影片」格式;台詞用 <d>[Chinese] …</d> 包起來)

可複製內容
subject_definitions:
<Picture 1> is the first frame of [Shot 1], showing <一句話描述你畫好的第一格>.
<Subject 1> is <主角> in <Picture 1>, with <三到五個外觀特徵>.

summary:
[keyframe completion + reference generation] The target video starts exactly on <Picture 1>. <一句話講整支影片發生什麼>.

retention_analysis:
<Picture 1> (appears in [Shot 1]): fully_preserved - used unchanged as the first frame, keeping <畫風>.
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - <要保留的外觀特徵> are retained.

detailed_description:
The target video is <畫風一句話,每一支都逐字重貼同一句>.
[Shot 1] The video opens exactly on <Picture 1>. <這一鏡唯一的動作>.
[Shot 2] At 00:03.500, the shot cuts to <景別> of <Subject 1> (S1). It says in <聲音描述> with a Taiwanese Mandarin accent, <d>[Chinese] <台詞>。</d> <說完之後的一個動作>.

overall_soundscape:
<環境音與動作音>

non_diegetic_music:
<配樂一句話,沒有就寫 N/A>

5. 踩坑排除

狀況原因怎麼辦
RuntimeError: Connection was lost.,指令卡住不退出一條長連線被網路中斷改用第 2 節的背景跑法;卡住的本機程序直接砍掉,再 colab stop
開機後第一次 exec 就斷線開機後先上傳、才第一次 exec先暖身 exec 再上傳
Allocation refused (precondition failed)可能已經有一台 GPU 機器在跑,或暫時沒有 GPU先 colab sessions 看有沒有忘了關的機器;這個錯誤不扣點
等前一個工作結束的迴圈一直不結束用 pgrep -f "xxx" 等程序,結果抓到監看指令自己改成檢查 colab sessions 裡的機器名稱
模型下載很慢、出現 HF_TOKEN 警告沒登入 Hugging Face,會被限速可以忽略;實測 55 GB 不到 4 分鐘就下載完
影片模型選到很大的 BF16 權重Colab 的 PyTorch 是 CUDA 12.8,A100 用不到 INT8 版正常;要更快可以自行升級 PyTorch 到 CUDA 13 再試

6. 除了影片還能跑什麼

同一套背景跑法可以跑任何 Python 工作。適合 Colab 的,都是「量大、要 GPU、跑完就收工」的事:

  • 一次轉完幾十小時的會議或課程錄音(Whisper 類模型)
  • 大量掃描文件轉文字(OCR 模型)
  • 用自己的資料訓練一個小的圖像風格模型或分類模型

實測環境:Colab CLI 0.7.4、macOS、Google AI Pro(台灣)、2026-09-28。Colab 的費率和規則會變,開機後先跑 colab --auth=oauth2 usage 看當下每小時扣幾點。

還有加碼

訂閱電子報,加碼拿我們實際用的批次工具,和三頭龍、三集短劇的完整提示詞。

送出即表示同意我們用這個信箱寄送電子報。我們只會用它寄信,不會給第三方。 你可以隨時退訂或要求刪除,詳見隱私權政策。