Map content & metadata fields, chunk, dedupe, and export ready-to-run Python/JS/TS code & vector-DB snippets.
Drag & drop a file here, or paste directly into the box below
Convert structured or semi-structured data into LangChain-style documents without writing a custom parser. Upload or paste JSON, CSV, Excel, Markdown, HTML, XML or YAML, select which fields should become document content, map useful fields to metadata, optionally split long content into chunks, preview the result and export it for a Python or JavaScript LangChain workflow.
This converter is designed for preparing data for Retrieval-Augmented Generation (RAG), semantic search, knowledge bases, document assistants and vector-store ingestion. Conversion prepares document objects only; it does not create embeddings, call a language model or upload the result to a vector database.
In LangChain Python, a Document represents a unit of text and associated information. Its main fields are:
page_content: The text that will be read, split, embedded or retrieved.metadata: A dictionary of contextual values such as source, URL, category, author, date or record ID.id: An optional string identifier supported by the current Python document abstraction.In LangChain JavaScript, the corresponding content property is commonly named pageContent. Keeping text separate from metadata lets a retrieval system search the content while filtering or citing results through structured fields.
page_content. Combine multiple fields with an appropriate separator when needed.| Format | Typical Source | Recommended Mapping |
|---|---|---|
| JSON | API responses and structured datasets | Text properties as content; IDs and categories as metadata |
| CSV / Excel | Catalogs, reports and tabular exports | One document per row, using selected columns |
| Markdown | README files and documentation | Readable body as content; filename or title as metadata |
| HTML | Articles and exported web content | Useful text as content; URL or title as metadata |
| XML | Feeds and enterprise records | Selected element text as content and attributes as metadata |
| YAML | Structured configuration or records | Human-readable fields as content and keys as metadata |
Suppose a product catalog contains these rows:
name,category,price,description
Trail Runner 200,Footwear,129.99,Lightweight trail running shoe
Pack Lite 30,Bags,89.50,30-liter hiking backpack
Select name and description as content fields. Map category and price to metadata. The first generated record can look like:
{
"page_content": "Trail Runner 200\nLightweight trail running shoe",
"metadata": {
"category": "Footwear",
"price": "129.99"
}
}
In a current Python LangChain project, the equivalent object is created with:
from langchain_core.documents import Document
document = Document(
page_content="Trail Runner 200\nLightweight trail running shoe",
metadata={"category": "Footwear", "price": "129.99"},
)
Place text that should influence semantic retrieval in page_content. Product names, descriptions, article paragraphs, support answers and documentation are good candidates. Put values used for filtering, organization or source attribution in metadata. Avoid using long descriptive text only as metadata because many retrieval workflows embed document content rather than every metadata value.
Chunk long articles, documentation or combined records when retrieving the whole document would return too much unrelated context. Smaller chunks can improve retrieval precision, while larger chunks preserve more surrounding meaning. There is no universal chunk size: test the result with your embedding model, retriever and query patterns. Avoid chunking already-short rows merely to create more documents.
page_content: Select at least one non-empty descriptive field.Document from langchain_core.documents.Conversion is described as browser-based, but you should still inspect the page and your browser’s network activity before processing confidential data. The tool transforms and maps source content; it does not confirm semantic quality, estimate embedding cost, validate vector-database compatibility or guarantee improved RAG answers. Final results depend on data quality, chunking, embeddings, retrieval settings and the language model.
page_content, associated metadata, and optionally an ID in current LangChain Python. It is commonly used by loaders, text splitters, retrievers and vector stores.