How to Convert a Word Document to Hierarchical JSON with Python
Question details
The user needs a programmatic way to convert a Word document into a JSON file while retaining the hierarchical parent-child relationships of document headings and nested lists.

- Product
- Microsoft Word / Python
- Device & OS
- not provided
- Scenario
- Automating the extraction of structured content from Word documents into a machine-readable JSON format for data processing, web rendering, or API integration.
- Observed behavior
- The user seeks to accurately map document heading levels and list indentations into nested JSON objects rather than extracting flat text.
Ensure you have Python installed on your system. You will also need to install the required library by opening your terminal or command prompt and running: pip install python-docx.
Use the python-docx Library to Parse and Serialize Document Structure
Extract headings and lists by identifying paragraph styles, map them into a nested dictionary based on hierarchy, and serialize the result using Python's built-in json module.
The python-docx library allows you to read paragraph text and inspect document styles. By checking if a paragraph style starts with 'Heading' or 'List', you can determine the structural role of each element.
To preserve the parent-child relationship, you must maintain a reference to the current active heading level as you iterate through the paragraphs, appending lower-level items as children.
In your Python script, import the necessary modules by writing 'from docx import Document' and 'import json'.
Instantiate the document object by pointing it to your file path: 'doc = Document("your_word_document.docx")'.
Create a loop iterating through 'doc.paragraphs'. Inside the loop, read 'paragraph.text' and check 'paragraph.style.name'. Use string matching to detect 'Heading 1', 'Heading 2', or 'List Paragraph'.
Create a state machine or stack in your code. When a Heading 1 is encountered, create a new dictionary object. When a Heading 2 or List item is found, append it to a 'children' array inside the current Heading 1 object.
Once the document iteration is complete and your dictionary is populated, use 'json.dumps(structure, indent=4)' to convert the Python dictionary into a formatted JSON string, which can then be written to a .json file.

Prepare Word Documents for Python Processing with WPS Office
While Python handles the automation, ensuring your documents have clean, consistent heading and list styles is crucial for accurate script parsing. WPS Office provides a free, lightweight, and highly compatible environment to format your .docx files seamlessly before running your Python scripts.
- 1. Open Your Document in WPS Writer: Launch WPS Office and open the .docx file you intend to process with your Python script.
- 2. Apply Standard Styles: Navigate to the Home tab. Select your text and use the Styles gallery to apply consistent 'Heading 1', 'Heading 2', and standard Bulleted Lists to ensure python-docx reads them correctly.
- 3. Save the Document: Click File > Save to retain the clean XML structure required for accurate automated JSON extraction.

Frequently Asked Questions
Why is python-docx not reading my nested list indentation correctly?
The python-docx library reads list items as standard paragraphs applied with a list style. It does not inherently return a tree structure for nested lists. To capture indentation, you must inspect the paragraph's XML attributes or its left indent property, and write custom Python logic to group them accordingly.
Can I extract tables into JSON using this method?
Yes, but tables are stored separately from paragraphs in the docx object model. You can access them via 'doc.tables'. To maintain the exact chronological order of tables alongside text paragraphs, you will need to iterate over the document's block-level elements using deeper XML parsing (python-docx's inner iter_block_items).
Does this Python conversion method work with older .doc files?
No, the python-docx library only supports the newer XML-based .docx format. To process older .doc files, you must first convert them to .docx. You can easily do this by opening the .doc file in WPS Office or Microsoft Word and saving it in the .docx format.




