AI system · metallurgy, regulatory documentation
GOST graph and RAG on steel requirements
EVRAZ · 2025 · AI system · Python 3.11, LangChain
From a set of PDF standards to a linked knowledge base: links between GOSTs, search with sources, export of steel class requirements to JSON
A series of prototypes for EVRAZ, whose GOST library is a set of PDFs with no links between documents. Standards (PDF/DOC) are converted to Markdown with tables and images, references to other GOSTs (normative and informative) are found automatically, “child” standards are downloaded, and RAG search with sources and structured data extraction are built on top. The test task is strength class C235 under GOST 27772-2021: the chemical composition and mechanical properties are exported to JSON. Over one week (30 October – 6 November 2025), five implementation options were tried in parallel.
The task
References to other GOSTs inside the documents are not clickable. To gather all the requirements for C235 steel, a specialist has to find the main standard by hand and then go through every standard it mentions one by one, which takes a long time and lets things slip through.
What’s inside
- Conversion of GOSTs from DOC/DOCX/PDF to Markdown, tables (XLSX) and images
- Automatic extraction of references classified as “normative / informative”; in GOST 27772-2021, 95 unique references were found (71 normative, 24 informative) and 232 clickable hyperlinks were inserted
- Automatic search and download of “child” GOSTs (allgosts.ru, with the Yandex XML API as a fallback): the corpus contains 60 .doc documents
- RAG answers with sources (Claude 4.5 Sonnet / Gemini 2.5 Pro via OpenRouter, FAISS vector index)
- Extraction of structured data for strength class C235 (8 chemical composition elements, 3 mechanical properties) to JSON with validation
- Variants: a web knowledge base on React/tRPC, a standards graph in Neo4j with graph-aware answers, RAGFlow + OpenSearch with k-NN without Docker, a FastAPI + HTMX web service with a crawler, a specification for a high-accuracy OCR module
Machine translation — being proofread.