GraphGen

GraphGen

Knowledge-driven synthetic data for LLM fine-tuning

Description

Domain fine-tuning needs quality QA data that is slow and costly to label by hand. GraphGen builds a knowledge graph from your corpus, finds what the model knows poorly and synthesizes targeted QA data.

Output plugs straight into LLaMA-Factory, XTuner and other SFT frameworks, with a full cookbook.

Features



Knowledge graph:Entities and relations.

Weak spots:Targets model gaps.

QA types:Atomic, aggregated and multi-hop.

Training ready:Mainstream frameworks.