Intelligent systems draw much of their reliability from the quality of their ontologies; however, manual ontology assessment remains patchy, time-consuming, and difficult to scale. To address these limitations, this paper proposes a domain-independent, machine-learningdriven framework for ontology quality assessment and improvement in the Semantic Web. The framework combines structural, semantic, and documentation metrics with supervised learning models to predict quality issues and recommend targeted refinements through a four-phase workflow comprising ML model development, metric definition, automated improvement, and empirical evaluation. The approach is validated on educational knowledge graphs using 1500 ontology modules from the EDUKG repository, including a 100-module expert-annotated gold set ( κ = 0.82). Experimental results show structural precision of 93.5% and semantic precision of 90.2%, with overall F1-scores close to 90%, while reducing ontology development time by 42% and quality assessment time by 65%.
| Authors | Jaziri, Wassim and Sassi, Najla |
|---|---|
| Year | 2026 |
| Venue | Systems |
| DOI | 10.3390/systems14020154 |
| Source Database | EBSCO (Applied Science & Technology) |
| Bridge-to-GNN | Category C |
| GNN Architecture | GNN (heuristic — verify) |
| Graph Encoding | Knowledge Graph / Ontological Network (inferred from title) |
| AEC Task | Knowledge Graph Construction (inferred from title) |
| Cohort | Mature Applications (2024–2026) |
| Dataset Size | 25 models (heuristic — verify) |
|---|---|
| Implementation Framework | PyTorch, scikit-learn, OWL / SPARQL (heuristic — verify) |
| Key Hyperparameters | lr 0.05; 16 hidden units (heuristic — verify) |
| Primary Metric | Accuracy (heuristic — verify) |
| Primary Value | 87.4% (heuristic — verify) |
| Key Finding | domain-independent, machine-learningdriven framework for ontology quality assessment and improvement in the Semantic Web (heuristic — verify) |
| Quality Assessment | 4/6 · rigor: Medium · Code/Data: ✓Benchmark: ✓Baseline: ✓Reproducible: ✓ (heuristic — verify) |
| GNN architectures detected | Knowledge Graph Embedding (TransE/RotatE/DistMult) |
|---|---|
| Frameworks / libraries | PyTorch · scikit-learn · RDF / OWL / SPARQL |
| Benchmark datasets referenced | PubMed |
| Code repositories found in text | https://github.com/THU-KEG/ · https://github.com/THU-KEG/EDUKG |
| Hyperparameters (regex-detected) | learning_rate=0.05 · hidden_units=16 · dropout=0.2 |
| Dataset stats found | 25 models · 300 trees · 200 trees |
| Reported metrics + values | Accuracy: 87.4% · F1 Score: 88.9% · F1 Score: 90% · F1 Score: 89% · F1 Score: 92.6% · F1 Score: 91% · F1 Score: 89.5% · F1 Score: 85% |
| Quality (heuristic, 6-flag) | 4/6 · rigor: Medium Code/Data: ✓Benchmark: ✓Baseline: ✓CV/Split: —Ablation: —Reproducible: ✓ |
Part of the GML/GNN in AEC Systematic Review (PRISMA 2020) — 112 papers, 2020–2026