Skip to content
#

synthetic-data

Here are 1,465 public repositories matching this topic...

Synthadoc: An open-source LLM knowledge compilation engine that turns raw documents into structured, local-first wikis. A transparent, human-readable alternative to traditional RAG, which can be self-managed and self-improved without the use of any tools.

  • Updated Aug 15, 2026
  • Python

Verbalized Sampling, a training-free prompting strategy to mitigate mode collapse in LLMs by requesting responses with probabilities. Achieves 2-3x diversity improvement while maintaining quality. Model-agnostic framework with CLI/API for creative writing, synthetic data generation, and dialogue simulation.

  • Updated Jan 3, 2026
  • Python

Improve this page

Add a description, image, and links to the synthetic-data topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the synthetic-data topic, visit your repo's landing page and select "manage topics."

Learn more