Watch It? Track It.

검색도 정적 색인입니다. 약 7,000개의 작은 파일이 모든 언어로 된 모든 시리즈, 영화, 인물의 이름을 함께 담고 있습니다. 쿼리 하나에는 그중 한 파일만 필요합니다.

파일

  • https://api.watchittrackit.com/v1/search/manifest.json
  • https://api.watchittrackit.com/v1/search/{build}/tree.json
  • https://api.watchittrackit.com/v1/search/{build}/l/{hex}.json
  • https://api.watchittrackit.com/v1/search/{build}/s/{hex}.json
  • https://api.watchittrackit.com/v1/search/recent.json

쿼리 조회하기

  1. v1/search/manifest.json을 읽어 현재 빌드 id를 확인합니다.
  2. v1/search/{build}/tree.json을 읽습니다. 빌드별로 절대 바뀌지 않으므로 한 번만 가져오면 됩니다.
  3. 쿼리를 정규화합니다. 소문자로 바꾸고 악센트와 문장 부호를 제거합니다. 불용어가 아닌 가장 긴 단어로 조회합니다.
  4. 그 단어로 트리를 따라가 파일을 찾습니다. l/의 버킷이거나, 단어가 분할 접두사에서 정확히 끝나면 s/의 요약입니다. 파일 이름은 접두사 UTF-8 바이트의 16진수입니다.
  5. 쿼리의 모든 단어가 이름의 어떤 단어의 시작 부분과 일치하는 항목만 남깁니다. 쿼리 전체로 시작하는 이름을 앞에 두고, 그다음 인기순으로 정렬합니다.
  6. v1/search/recent.json을 적용합니다. 빌드 이후 변경된 레코드가 나열되어 있습니다. 해당 레코드의 색인 항목은 버리고 이 항목을 대신 사용합니다. 레코드가 수정될 때마다 바뀌므로, 새 사본을 받으려면 몇 분마다 바뀌는 쿼리 문자열(?w=에 현재 5분 구간 값)을 붙이세요.

항목

각 항목은 배열입니다:

위치유형설명
0string레코드 유형.
1number레코드 id.
2string일치한 이름(언어는 무엇이든 가능).
3number유형 내 순위: 1,000,000이 가장 인기가 많습니다.
4string레코드 슬러그.
5number | null연도(작품의 경우).
6string | null썸네일 경로.

예시(JavaScript)

const BASE = "https://api.watchittrackit.com/v1/search"
const STOPWORDS = new Set(["a", "an", "and", "the", "of", "in", "on", "to", "for", "at", "by", "with", "or",
  "de", "del", "la", "le", "les", "el", "los", "las", "lo", "un", "une", "et", "du", "des", "da", "do", "dos",
  "der", "die", "das", "und", "ein", "eine", "den", "dem", "von", "zu", "il", "i", "e", "di", "y", "en", "het", "een", "van"])

const normalize = (s) => s.normalize("NFKD").replace(/\p{M}/gu, "").toLowerCase()
  .replace(/[^\p{L}\p{N}]+/gu, " ").trim()
const hex = (s) => [...new TextEncoder().encode(s)].map((b) => b.toString(16).padStart(2, "0")).join("")

async function search(query) {
  const { build } = await fetch(`${BASE}/manifest.json`).then((r) => r.json())
  const tree = await fetch(`${BASE}/${build}/tree.json`).then((r) => r.json())

  const q = normalize(query)
  const words = q.split(" ").filter(Boolean)
  const real = words.filter((w) => !STOPWORDS.has(w))
  const word = (real.length ? real : words).sort((a, b) => [...b].length - [...a].length)[0]

  // Walk the tree: descend while the prefix is split; stop at a bucket or a summary.
  let key = "", file
  for (const c of word) {
    if (tree[key + c]) { key += c; continue }
    const starts = tree[key] ?? []
    const start = starts.filter((s) => s <= c).pop() ?? starts[0] ?? c
    file = `l/${hex(key + start)}.json`
    break
  }
  file ??= `s/${hex(key)}.json`

  const res = await fetch(`${BASE}/${build}/${file}`)
  let entries = res.ok ? await res.json() : []

  // Records changed since the build replace the index's entries for them.
  // The query string changes every 5 minutes, so caches don't serve an old copy.
  const recent = await fetch(`${BASE}/recent.json?w=${Math.floor(Date.now() / 300000)}`).then((r) => r.json())
  entries = entries.filter(([type, id]) => !(`${type}:${id}` in recent.records))
  entries.push(...Object.values(recent.records).flat())

  const seen = new Set()
  return entries
    .filter((e) => { const n = normalize(e[2]).split(" "); return words.every((w) => n.some((x) => x.startsWith(w))) })
    .sort((a, b) => Number(normalize(b[2]).startsWith(q)) - Number(normalize(a[2]).startsWith(q)) || b[3] - a[3])
    .filter(([type, id]) => !seen.has(`${type}:${id}`) && seen.add(`${type}:${id}`))
    .slice(0, 30)
}

// [["tv", 81189, "Breaking Bad", 999998, "breaking-bad", 2008, "/images/tv/81189/….t.jpg"], …]
console.log(await search("breaking bad"))

색인은 매일 다시 빌드됩니다. 빌드의 파일은 절대 바뀌지 않으므로 원하는 만큼 캐시해도 됩니다. 바뀌는 것은 manifest와 recent.json뿐입니다.