Watch It? Track It.

搜索同样是静态索引:约 7,000 个小文件共同涵盖所有语言下每部剧集、电影和每个人物的所有名称。一次查询只需其中一个文件。

文件

  • https://api.watchittrackit.com/v1/search/manifest.json
  • https://api.watchittrackit.com/v1/search/{build}/tree.json
  • https://api.watchittrackit.com/v1/search/{build}/l/{hex}.json
  • https://api.watchittrackit.com/v1/search/{build}/s/{hex}.json
  • https://api.watchittrackit.com/v1/search/recent.json

执行查询

  1. 读取 v1/search/manifest.json 获取当前构建 id。
  2. 读取 v1/search/{build}/tree.json。同一构建中它永远不会改变,因此只需获取一次。
  3. 规范化查询:转为小写,去除重音符号和标点。用其中不是停用词的最长单词进行查找。
  4. 用该单词遍历树以找到对应文件:l/ 中的一个桶,或者当单词恰好结束于某个拆分前缀时,s/ 中的一个摘要。文件名是前缀 UTF-8 字节的十六进制表示。
  5. 保留那些查询中每个单词都是名称中某个单词开头的条目。以完整查询开头的名称排在最前,其余按热度排序。
  6. 应用 v1/search/recent.json:它列出了自构建以来发生变化的记录。丢弃索引中这些记录的条目,改用此文件中的条目。它会随记录编辑而变化,因此请附加一个每隔几分钟变化的查询字符串(?w= 加上当前的 5 分钟时间窗口)以获取最新副本。

条目

每个条目是一个数组:

位置类型说明
0string记录类型。
1number记录 id。
2string匹配到的名称,可以是任意语言。
3number在其类型内的排名:1,000,000 为最热门。
4string记录 slug。
5number | null年份(适用于作品)。
6string | null缩略图路径。

示例(JavaScript)

const BASE = "https://api.watchittrackit.com/v1/search"
const STOPWORDS = new Set(["a", "an", "and", "the", "of", "in", "on", "to", "for", "at", "by", "with", "or",
  "de", "del", "la", "le", "les", "el", "los", "las", "lo", "un", "une", "et", "du", "des", "da", "do", "dos",
  "der", "die", "das", "und", "ein", "eine", "den", "dem", "von", "zu", "il", "i", "e", "di", "y", "en", "het", "een", "van"])

const normalize = (s) => s.normalize("NFKD").replace(/\p{M}/gu, "").toLowerCase()
  .replace(/[^\p{L}\p{N}]+/gu, " ").trim()
const hex = (s) => [...new TextEncoder().encode(s)].map((b) => b.toString(16).padStart(2, "0")).join("")

async function search(query) {
  const { build } = await fetch(`${BASE}/manifest.json`).then((r) => r.json())
  const tree = await fetch(`${BASE}/${build}/tree.json`).then((r) => r.json())

  const q = normalize(query)
  const words = q.split(" ").filter(Boolean)
  const real = words.filter((w) => !STOPWORDS.has(w))
  const word = (real.length ? real : words).sort((a, b) => [...b].length - [...a].length)[0]

  // Walk the tree: descend while the prefix is split; stop at a bucket or a summary.
  let key = "", file
  for (const c of word) {
    if (tree[key + c]) { key += c; continue }
    const starts = tree[key] ?? []
    const start = starts.filter((s) => s <= c).pop() ?? starts[0] ?? c
    file = `l/${hex(key + start)}.json`
    break
  }
  file ??= `s/${hex(key)}.json`

  const res = await fetch(`${BASE}/${build}/${file}`)
  let entries = res.ok ? await res.json() : []

  // Records changed since the build replace the index's entries for them.
  // The query string changes every 5 minutes, so caches don't serve an old copy.
  const recent = await fetch(`${BASE}/recent.json?w=${Math.floor(Date.now() / 300000)}`).then((r) => r.json())
  entries = entries.filter(([type, id]) => !(`${type}:${id}` in recent.records))
  entries.push(...Object.values(recent.records).flat())

  const seen = new Set()
  return entries
    .filter((e) => { const n = normalize(e[2]).split(" "); return words.every((w) => n.some((x) => x.startsWith(w))) })
    .sort((a, b) => Number(normalize(b[2]).startsWith(q)) - Number(normalize(a[2]).startsWith(q)) || b[3] - a[3])
    .filter(([type, id]) => !seen.has(`${type}:${id}`) && seen.add(`${type}:${id}`))
    .slice(0, 30)
}

// [["tv", 81189, "Breaking Bad", 999998, "breaking-bad", 2008, "/images/tv/81189/….t.jpg"], …]
console.log(await search("breaking bad"))

索引每天重建。构建的文件永远不会改变,因此可以任意长时间缓存;只有 manifest 和 recent.json 会变化。