Neo4j Document Intelligence
Neo4j Document Intelligence is a document intelligence feature based on Aura, released by Neo4j in June 2026. It converts PDF, DOCX, TXT, EPUB, HTML, and Markdown documents into searchable and queryable knowledge graphs. Researchers can provide documents and describe their desired analysis objectives in natural language, and the system will suggest a graph schema and use large language models to extract entities and relationships. If a document repository is simply a bookshelf for storing files, this tool is more like an interactive map that connects the concepts and evidence within the documents. The result allows for full-text search.
Neo4j Document Intelligence is a document intelligence feature based on Aura, released by Neo4j in June 2026. It converts PDF, DOCX, TXT, EPUB, HTML, and Markdown documents into searchable and queryable knowledge graphs. Researchers can provide documents and describe their desired analysis objectives in natural language, and the system will propose a graph schema and extract entities and relationships using large language models. If a document repository is simply a bookshelf for storing files, this tool is more like an exploratory map that connects concepts and evidence within the documents. The results can be used in a GraphRAG configuration, which utilizes both a lexical graph for original text search and an entity graph to represent semantic relationships.
In traditional document-based knowledge graph construction, file parsing, chunk splitting, named entity recognition, relationship extraction, schema design, and graph loading all had to be implemented separately. In particular, research fields differ in their entity types and relationships, which led to problems where the initial graph model did not match actual questions, or where multi-step relationships scattered throughout the documents could not be explained by vector search alone. The key difference of Neo4j Document Intelligence is that it creates schema candidates from natural language instructions and constructs the relationship structure of the document within Aura, without requiring separate extraction code or prior graph modeling. Just as GPT generates answers based on the meaning of the text, this tool transforms the entities and relationships hidden in the text into nodes and links in the graph, creating a foundation for subsequent exploration.
A life science researcher can input research papers in PDF format and Markdown research notes, and then request a schema in natural language that includes entities such as Gene, Protein, Disease, and Drug, and relationships such as INTERACTS_WITH, ASSOCIATED_WITH, and TARGETS. In the generated entity graph, it is possible to trace multi-step paths to see how a specific drug is connected to proteins and diseases in multiple papers, and in the lexical graph, it is possible to search for relevant sections of the original text to confirm the evidence for the relationship. This is useful for organizing connections between documents that might be missed by simple keyword searches into candidate hypotheses or relationships to be reviewed.
In addition, clinical and regulatory documents can be processed together, including DOCX protocols, PDF reports, and HTML guidelines, to structure the relationships between trial stages, evaluation variables, adverse events, and target diseases. The GraphRAG application can be configured to narrow down relevant paths in the entity graph and then search for the original text context in the lexical graph. However, since LLM-based extraction results do not automatically guarantee the facts in the original text, it is necessary to verify the sources for each relationship and have expert review before using them in life science and clinical decision-making. Supported limits, extraction model selection, data retention, and security conditions should be further verified in the official Aura documentation.
💻 System Requirements
Local RAM requirements need to be verified based on Neo4j Aura managed service specifications
Local VRAM/GPU requirements need to be verified
Document count, graph scale, and Aura plan limits need to be confirmed
⚡ Installation
4-1. Quick Start
The official installation command was not confirmed from the provided Discovery information. Since it is introduced as a managed service using the Document Intelligence feature of Neo4j Aura, you must verify the account, plan, and feature activation procedures in the official documentation.
4-2. Detailed installation
No arbitrary commands are listed because verified CLI, SDK, or API examples from official documentation have not been secured. Installation and initial setup require consulting the latest guidance for https://neo4j.com/docs/aura/document-intelligence/introduction/.
🧬 Bio Use Cases
🔬 Literature-Based Drug-Target Knowledge Graph
Input research paper PDFs and Markdown research notes, and request Drug, Protein, Disease entities and TARGETS, ASSOCIATED_WITH relationships. Query the multi-hop paths of the entity graph and the original text segments of the lexical graph together to identify candidate drugs for repurposing.
🧬 Gene-Phenotype Evidence Integration
Combine PDF research papers, HTML data descriptions, and TXT notes to obtain a Gene, Variant, Phenotype, Study schema. Integrate relationships from different documents into a graph and review the original source evidence to support disease-specific variant interpretation and prioritize subsequent experiments.
📋 Clinical and Regulatory Document Relationship Exploration
Process DOCX clinical protocols and PDF reports to extract Trial, Endpoint, AdverseEvent, Intervention relationships. Use GraphRAG queries to explore the connections between evaluation variables and adverse events across trials, and verify the original context to narrow down the scope of document review.
FAQ
What is Neo4j Document Intelligence?
Neo4j Document Intelligence is a document intelligence feature based on Aura, released by Neo4j in June 2026. It converts PDF, DOCX, TXT, EPUB, HTML, and Markdown documents into searchable and queryable knowledge graphs. Researchers can provide documents and describe their desired analysis objectives in natural language, and the system will propose a graph schema and extract entities and relationships using large language models. If a document repository is simply a bookshelf for storing files, this tool is more like an exploratory map that connects concepts and evidence within the documents. The results can be used in a GraphRAG configuration, which utilizes both a lexical graph for original text search and an entity graph to represent semantic relationships. In traditional document-based knowledge graph construction, file parsing, chunk splitting, named entity recognition, relationship extraction, schema design, and graph loading all had to be implemented separately. In particular, research fields differ in their entity types and relationships, which led to problems where the initial graph model did not match actual questions, or where multi-step relationships scattered throughout the documents could not be explained by vector search alone. The key difference of Neo4j Document Intelligence is that it creates schema candidates from natural language instructions and constructs the relationship structure of the document within Aura, without requiring separate extraction code or prior graph modeling. Just as GPT generates answers based on the meaning of the text, this tool transforms the entities and relationships hidden in the text into nodes and links in the graph, creating a foundation for subsequent exploration. A life science researcher can input research papers in PDF format and Markdown research notes, and then request a schema in natural language that includes entities such as Gene, Protein, Disease, and Drug, and relationships such as INTERACTSWITH, ASSOCIATEDWITH, and TARGETS. In the generated entity graph, it is possible to trace multi-step paths to see how a specific drug is connected to proteins and diseases in multiple papers, and in the lexical graph, it is possible to search for relevant sections of the original text to confirm the evidence for the relationship. This is useful for organizing connections between documents that might be missed by simple keyword searches into candidate hypotheses or relationships to be reviewed. In addition, clinical and regulatory documents can be processed together, including DOCX protocols, PDF reports, and HTML guidelines, to structure the relationships between trial stages, evaluation variables, adverse events, and target diseases. The GraphRAG application can be configured to narrow down relevant paths in the entity graph and then search for the original text context in the lexical graph. However, since LLM-based extraction results do not automatically guarantee the facts in the original text, it is necessary to verify the sources for each relationship and have expert review before using them in life science and clinical decision-making. Supported limits, extraction model selection, data retention, and security conditions should be further verified in the official Aura documentation.
When should I use Neo4j Document Intelligence?
Neo4j Document Intelligence is a document intelligence feature based on Aura, released by Neo4j in June 2026. It converts PDF, DOCX, TXT, EPUB, HTML, and Markdown documents into searchable and queryable knowledge graphs. Researchers can provide documents and describe their desired analysis objectives in natural language, and the system will suggest a graph schema and use large language models to extract entities and relationships. If a document repository is simply a bookshelf for storing files, this tool is more like an interactive map that connects the concepts and evidence within the documents. The result allows for full-text search.
What is a biomedical use case for Neo4j Document Intelligence?
🔬 Literature-Based Drug-Target Knowledge Graph: Input research paper PDFs and Markdown research notes, and request Drug, Protein, Disease entities and TARGETS, ASSOCIATED_WITH relationships. Query the multi-hop paths of the entity graph and the original text segments of the lexical graph together to identify candidate drugs for repurposing.
📝 Update Notes
No update notes yet.
🧪 Related Code of Life
No related Code of Life posts yet.