MadeInRed
GEO Skills

JSON-LD Knowledge Graph

在 llms.txt 旁边放一份 schema.org 知识图谱,方便模型按实体引用。

shimo4228/jsonld-knowledge-graph开发者、内容工程1 个已核验 Skill 文件

需要提供什么

稳定的概念结构

你会得到什么

graph.jsonld

JSON-LD Knowledge Graph

graph.jsonld を llms.txt の隣に置いて、project の concept-level architecture を schema.org JSON-LD triples として encode する skill。LLM が prose だけでは引き出しにくい entity 間の関係(matrix / hierarchy / phase-binding)を triple として citation 可能にする。

When to Use

以下を すべて 満たす project に適用する:

  • 安定した concept-level 構造を持つ
    • 例: matrix(4 Quadrants × N ADRs)、ordered hierarchy(prohibition-strength 3 levels)、phase-skill binding(6:6 bijective)、layered architecture(3 memory layers)
  • 構造が release を跨いで安定(vX.Y.Z bump で entity が増減しない)
  • LLM に当該関係を訊いた時に prose だけでは答が出にくい / 誤答する 経験的証拠がある
  • すでに llms.txt + llms-full.txt(Answer.AI 標準)を持っている

When NOT to Use

以下のいずれかに該当する場合は 適用しない:

  • 単一目的の linear project(matrix / hierarchy 構造がない)
  • 内部構造が release ごとに churn する(graph が編集負債になる)
  • llms-full.txt の prose Q&A で関係問題が既に十分答えられる
  • 「LLM citation 向上を狙いたい」だけで具体的な entity 関係問題がない(GEO 目的なら llms-txt-writer の方が direct)

graph.jsonld は prose の代替ではなく 補完。prose で表現しきれない構造があるとき、その構造をやっと拾える。

file-level 構造との関係(正本)

graph.jsonld は concept 層だけ を保存する。file-level 構造(「X はどのファイルに住むか」「誰が誰を呼ぶか」)は保存せず、問いのたびにコードから導出する:

file-level(導出、保存しない) graph.jsonld(保存)
対象 ファイル / モジュール / call 関係 概念 / エンティティ
答える質問 「X はどのファイルに住んでいるか」「誰が Y を呼ぶか」 「X とは何か、X と Y はどう関係するか」
源 コードそのもの — Claude Code の LSP tool(workspaceSymbol / findReferences / incomingCalls)、grimp 等の import グラフ JSON-LD triples(手で書く)
主読者 作業中の agent AI search engine + LLM が entity を citation する時
trigger なし(都度計算) concept / 関係の変化

graph にコードのノード(file path・module・LOC)を置かない。置くと第 2 の module map になり、ソース commit ごとに同期コストを払う鏡が育つ(手書き module map 6 枚がソース 197 commit に対し 159 commit の同期を要し、読者は観測されなかった実例がある)。設計理由は ADR、パイプラインの段構成はそれを走らせる script の冒頭コメントが持つ。

Drift 防止のための運用規約

  • 新規 ADR / Concept / Quadrant 等を追加 → graph.jsonld にノード追加(@id は GitHub blob URL)
  • ファイルパスが変わった → graph.jsonld の @id が blob URL ならそこだけ追従。他に更新する文書は無い
  • 概念の semantics が変わった → graph.jsonld の description 更新

context-sync skill(Maintain phase 担当)がこの drift を audit する。詳細は ~/.claude/skills/context-sync/SKILL.md 参照。

Required Design Moves

実装で観察された 9 つの再利用可能な move。すべて適用が default、外す場合は理由を明示する。

1. Dual @type on every node

{
  "@id": "...",
  "@type": ["ResearchLine", "ScholarlyArticle"]
}

custom namespace type(domain semantics を持つ)+ schema.org base type(AI search engine が citation target として認識)。両方つけることで、schema.org に親しい crawler も custom vocab に親しい LLM も同じノードを認識できる。

2. Language-tagged literals for bilingual

"alternateName": [
  {"@value": "Four Business AI Quadrants", "@language": "en"},
  {"@value": "ビジネスAIの四象限", "@language": "ja"}
]

1 file で en/ja 両対応。graph.ja.jsonld を別ファイルにするのは禁止(同期負債が発生する)。Concept DOI / repo URL / @id は言語中立なので二重管理不要。

  • ResearchLine の name / alternateName、および concept-level node(Concept / Axiom / Quadrant 等)の alternateName は両言語で置く
  • 多言語化する literal には必ず @language を付ける(無いと default 言語が不明で crawler の解釈が undefined)
  • description と ADR title は en 単独でよい(prose の正本は llms-full.txt / GitHub の英語 prose、graph は entity 識別が主目的)

3. Schema absence で禁止関係を構造的に強制

「X と Y は 絶対に dependency 関係ではない」を保証したい場合、dependsOn のような edge type を そもそも @context に定義しない。schema の vocabulary 自体が構造的 commitment になり、表現できない関係は表現されない。

例: 3 つの sibling research line が独立進化することを強制するため、共通 vocab には dependsOn edge type を そもそも定義しない。代わりに siblingOf(symmetric)と derivesFrom(historical / 一方向、ある line が別 line から実装的に派生した事実を encode する場合のみ使う)の 2 種類だけを定義する。

4. Cross-graph @id reuse で triple merge

複数の graph.jsonld ファイルが存在する場合(hub-and-spoke topology)、共有 entity の @id は byte-identical にする。例:

  • 3 つの memory layers が 2 つの sibling line repo の graph に登場
  • 両方とも @id: "https://<owner>.github.io/<owner>/vocab#memory-layer/episode-log" を使う(owner namespace は project ごとに置換)
  • LLM がふたつの graph を crawl すると triple が自動 merge され、「同じ概念」として認識される

これは JSON-LD 1.1 の標準動作。pyld など conformant processor で実装可能。

5. Volatile state を schema レベルで除外

graph.jsonld に 以下の field を持たせない:

  • version 番号 (version, versionNumber, vX.Y.Z)
  • count(ADR 数、test 数、ノード数)
  • 流動的 enumeration(churning skill list, dynamic capability list)

schema に存在しない field は entity に乗らないので、routine release が graph に漏れない。同梱 lint(Verification Workflow)の VOLATILE 検査は key "version" / "versionNumber" / "adrCount" / "testCount" の出現を検出する。それ以外の count field・vX.Y.Z 形の値・流動的 enumeration は lint が見ないので、作成時と review で目で確認する。

6. Matrix encoding via paired edges

A × B matrix を encode する場合、片方向の edge だけでは不十分。2 つの complementary edge を使う:

  • appliesTo (ADR → Quadrant): 各 ADR が適用される Quadrant 群
  • realizedBy (ProhibitionLevel → ADR): 各 prohibition tier が実現される ADR

silence is signal: 「Quadrant 1, 2 (Script, Algorithmic Search) は AAP 適用外」は appliesTo edge の 不在 で表現する。governanceTier: "out-of-scope" プロパティで明示的にも stamp する。

7. Root Dataset node で graph self-description

@graph array の先頭に schema.org Dataset 型のルートノードを置く:

{
  "@id": "https://github.com/owner/repo#knowledge-graph",
  "@type": ["Dataset", "CreativeWork"],
  "name": "Project Knowledge Graph",
  "description": "...",
  "isBasedOn": "https://github.com/owner/repo",
  "mainEntity": [<primary entity @ids>]
}

schema.org Dataset は Google AI Overviews / Perplexity が structured data として認識する。graph の 目的と entry points を crawler に明示する役割。

8. Reading-order block で AI 入口導線を明示

llms.txt 冒頭に Graph-first reading order block を置く。詳細手順は llms-txt-writer skill 参照。要点:

  • タイトル直下に > AI agents should read graph.jsonld first blockquote
  • 直後に numbered "Recommended reading order" section
  • Core documentation navigator の最上位に graph.jsonld

9. Reverse-link で hub-and-spoke 経路双方向化

hub-and-spoke topology の場合、各 line repo の README に hub graph への逆リンクを置く。README 側の置き方の正本は readme-writer(末尾に平文 1–2 行)。LLM が個別 line repo から入ってきた場合でも hub graph へ戻れる。1 段目の探索後に broader context へ広げる経路。

散文リンクだけでは不十分 — README 逆リンクは人間/プレーンテキスト向け。グラフ層でも各 spoke の self-node に機械可読な上向き edge を張る:

// 各 line / supporting repo の self-node (ResearchLine or Dataset) に:
"isPartOf": "https://github.com/<owner>/<hub>"   // hub の canonical @id (bare-repo URL に統一)
  • hub→spoke は mainEntity / siblingOf、spoke→hub は isPartOf で双方向化 (isPartOf の厳密な逆は hasPart だが、独立 triple として有効)
  • target @id は 全 repo で同一の canonical hub node に揃える (bare-repo github.com/<owner>/<hub> 推奨)。揃わないと merge 時に hub が複数ノードに割れて join 不能
  • isPartOf は context に {"@id":"https://schema.org/isPartOf","@type":"@id"} を宣言 (無いと「IRI 文字列の literal 化」と同じく literal になる)。context を触れない content graph では inline {"@id":"..."} で node 参照を強制
  • self-node を持たない content-only graph (Article 列挙のみ等) には最小 Dataset self-node を1個新設し、そこに isPartOf を張る
  • 監査: 新 repo 追加時、全 spoke の self-node が isPartOf→hub を持つか確認。実地では下向き (hub→spoke) だけ張られ、上向きが大半の spoke で抜けていた

Schema Vocabulary 設計

新規 type / edge を導入する判断基準:

新規 node type を作る条件

  • closed / bounded set: 「4 axioms」「3 memory layers」のように member が limited で stable
  • ordering invariant が必要: ProhibitionLevel のように level integer で順序を保つ必要がある
  • distinct semantics: Axiom は swappable preset の 1 instance、Concept は free-form term — 同じ Concept 扱いだと意味が混ざる

これに当てはまらない場合は、既存の Concept instance として encode する(4 code-LLM patterns のように、4 件しかなく新規 type 化に値しない場合)。

新規 edge type を作る条件

  • 既存 edge では semantic が違う: extends (EcosystemRepo → ResearchLine) と appliesTo (ADR → Quadrant) は別の関係。reuse すると混乱
  • matrix の片側を成す: appliesTo + realizedBy のように、matrix encoding に必須

新規 edge は schema.org base vocabulary に該当があれば優先する(url, identifier, citation, isBasedOn, mainEntity 等)。custom vocab に追加するのは schema.org base で表現できない時のみ。

@id namespace 設計

  • ResearchLine: 既存 DOI URL を使う(https://doi.org/10.5281/zenodo.NNNN、concept DOI 必須、latest DOI は使わない)
  • EcosystemRepo: GitHub repo URL(https://github.com/owner/repo)
  • Concept: custom vocab namespace(https://<owner>.github.io/<owner>/vocab#concept/<slug> 等)
  • ADR: GitHub blob URL(https://github.com/owner/repo/blob/main/docs/adr/NNNN-slug.md)

namespace は dereferenceable である必要なし(典型的な "private vocab" pattern)。重要なのは byte-identical re-use が cross-graph で実現できること。

sameAs の解決先は self-sovereign または earned なもののみ(ORCID / DOI / 自アカウントの platform profile / 無関係な第三者が作成した record)。Wikidata QID は張らない(host governance による一括削除を実測した)。dead QID を検出したら purge する。

Cross-graph @id Discipline

hub-and-spoke topology で複数 graph を持つ場合の規約:

  1. Hub graph に登場する entity の @id を canonical とする
  2. 各 line graph で同じ entity を参照する時は、copy-paste で @id を再利用(手入力しない、typo 防止)
  3. 新規 entity を追加する時は hub graph 側にまず登録してから line graph で参照
  4. @id を変更する場合は、すべての graph を同時に更新(cross-graph grep で漏れ確認)

verification:

# Hub graph の primary entity @id を抽出
jq -r '.["@graph"][] | select(.["@type"][] | contains("ResearchLine")) | .["@id"]' hub/graph.jsonld

# 各 line graph でも同じ @id が登場することを確認
grep -h "doi.org/10.5281/zenodo" */graph.jsonld | sort -u

Companion File Wiring

graph.jsonld を作っただけでは crawler に見つからない。llms.txt / llms-full.txt / README に wiring が必要:

ファイル 何を追加
llms.txt (1) 冒頭に Graph-first reading order blockquote と numbered section、(2) Core documentation navigator の最上位に graph.jsonld entry
llms-full.txt 末尾に question-form H2("How do X and Y relate as a graph?")+ graph.jsonld への link + 3 つ程度の load-bearing design choice 説明
README.md と維持している language mirror 末尾に平文 1–2 行で graph.jsonld / llms.txt への導線(hub-and-spoke の line README はこの行に hub graph への reverse-link も含める)。置き方の正本は readme-writer

詳細な wording は ~/.claude/skills/llms-txt-writer/SKILL.md の Companion JSON-LD Graph セクション参照。

Verification Workflow

graph.jsonld を作成・編集したら、同梱の lint script を走らせる。JSON validity / expansion / DROPPED-KEY / URL-LITERAL / VOLATILE を一括で決定論的に検査する:

# exit 0 = clean, 1 = findings, 2 = fatal。複数ファイル可
uv run --with pyld python3 ~/.claude/skills/jsonld-knowledge-graph/scripts/graph_lint.py graph.jsonld

# locator 系 property(url / license / contentUrl)は literal URL を許容。変更する場合:
#   --allow-literal-url url,license,contentUrl,codeRepository

# 複数 graph を渡すと cross-file の NAME-DRIFT も検査する
# (同一 IRI に複数の異なる en name → 検出は構造的・決定論的。どれを正とするかの
#   解決だけが judgment なので、lint は報告のみ。表記揺れの正解は alternateName に置く)
uv run --with pyld python3 ~/.claude/skills/jsonld-knowledge-graph/scripts/graph_lint.py hub/graph.jsonld line1/graph.jsonld line2/graph.jsonld
#   --skip-name-drift で省略可

修正の定石: 値の書き換えではなく @context への coercion 追加 で直す("sameAs": {"@id": "https://schema.org/sameAs", "@type": "@id"} を足せば既存の文字列値がそのまま IRI として解釈される。diff が context の数行で済む)。

Context pitfalls — edge が静かに消える 2 パターン

production graph 6 本の監査 (2026-06) で実際に発生した。いずれも JSON は valid のまま triple count も変わらないので、構文チェックと expansion だけでは見えない(lint の DROPPED-KEY / URL-LITERAL が拾う層):

  1. IRI 文字列の literal 化: @vocab があっても、context で @type: "@id" 強制の ない property(schema.org の author / creator 等)に IRI を文字列で渡すと literal になり、node への edge にならない。Person node が graph に居ても誰からも参照されない。
    • 対策: IRI 参照は常に {"@id": "https://..."} オブジェクト形式で書く
  2. 未定義 key の無言ドロップ: @vocab なしの明示マッピング型 context では、context に 無い key(sameAs 等)が expansion で警告なく消える。
    • 対策: graph に新しい property を導入する前に context マッピングの存在を確認する

Manual checks

  • llms.txt reading order: head -30 llms.txt | grep -c "Recommended reading order\|graph.jsonld"
  • Concept DOI 整合: 各 ResearchLine @id を Zenodo で開き、latest ではなく parent record(concept DOI)であることを確認
  • JSON-LD playground: https://json-ld.org/playground/ に paste して @context が解決し triple として展開されることを確認
  • schema.org validator: https://validator.schema.org/ で Dataset / ScholarlyArticle 等の type が認識されることを確認
  • LLM citation probe: graph push 後 1-2 週間(crawler refresh 待ち)、ChatGPT / Perplexity に「 の X と Y はどう関係しますか?graph.jsonld を参照してください」と質問し、graph が citation されるか確認

Mirror Sync to Hugging Face Datasets

graph.jsonld を更新したら、Hugging Face Datasets 上の mirror にも同期する。HF は LLM training pipeline / knowledge-graph crawler の primary ingest source として機能する(HF dataset は Auto-converted to Parquet が走り、pandas / Polars / Datasets ライブラリから直接 load 可能になる)。

手順・前提条件・repo mapping・token scope の正本は hf-sync。release-doi skill の Phase 4 末尾(tag push + gh release create の後)で呼ぶのが標準フローで、ad-hoc resync にも同じ skill を使う。

Maintenance Contract

graph.jsonld を 編集する trigger:

  • ecosystem repo の追加 / 退役 → EcosystemRepo ノード add/remove
  • 新しい stable structural concept の登場 → Concept / Axiom / Quadrant ノード add + 関連 edges
  • 新 research line の開始 → ResearchLine ノード + siblingOf edges
  • Concept DOI 自体が移動(稀、Zenodo record restructuring 時のみ)

graph.jsonld を 編集してはいけない trigger(routine release では触らない):

  • 任意 module の vX.Y.Z release
  • ADR count, skill count, test count, version bump
  • 内部 module の restructuring(file-level は保存していないので触る文書が無い。blob URL の @id だけ追従)
  • model version bump(qwen3.5:9b から qwen4:9b 等)

schema に version / count / churning field を持たせていない限り、これらは graph に そもそも encode できない ので構造的に強制される。

Reference Implementation

このパターンの canonical implementation pointer(具体 repo URL / DOI / vocab namespace)は本 skill 本文に含めず、同一ディレクトリの inspiration.md に記録する(portability 確保)。pattern を新しい project に適用する場合、以下の構造目安が参考になる:

  • Hub-and-spoke 構成: 1 つの hub graph が sibling line graph 群を siblingOf で繋ぐ
  • Hub graph 規模目安: 3-5 ResearchLine + 10-20 EcosystemRepo + 5-10 Concept + 3-5 ExternalReference
  • Per-line graph 規模目安: 1 ResearchLine + line 内部構造(matrix / hierarchy / phase-binding)を 20-40 nodes で encode
  • 共通 vocab: 全 graph が同じ namespace を使い、shared @id で triple merge を実現

単一 repo(hub なし)でも matrix / hierarchy 構造があれば適用可能。その場合 hub-related の sibling/derivesFrom 系 edge は不要。

What This Skill Does NOT Do

  • llms.txt / llms-full.txt の文章設計 — use llms-txt-writer
  • Project doc role の overlap 検出 / 整理 — use context-sync
  • file-level の module map の生成 — 作らない(LSP tool / grimp で都度導出。上の「file-level 構造との関係」)
  • Articles / blog post の文体 設計 — use writing-ecosystem(~/MyAI_Lab/zenn-content 常駐)

Related

  • llms-txt-writer — llms.txt / llms-full.txt 本体の書き方、navigator 設計、GEO/AEO 最適化
  • context-sync — graph.jsonld と prose docs の drift audit(Maintain phase)