Using Artificial Intelligence and Language Models to Make Occupational Risk Prevention More Predictive


Using Artificial Intelligence and Language Models to Make Occupational Risk Prevention More Predictive

Using Artificial Intelligence and Language Models to Make Occupational Risk Prevention More Predictive


Feature by EBIS-HSE | Tue 22nd Sep 2026

Summary

Artificial intelligence can help organisations analyse accident reports, near-miss descriptions, inspection findings and other safety records that are difficult to process consistently at scale.

This review examined 123 studies published between 2013 and October 2025 across construction, aviation, mining, chemical and process industries, transport, healthcare and other occupational settings. It considered natural language processing, machine learning, large language models and multimodal systems.

Earlier applications concentrated on classifying incidents and identifying recurring causes. Newer systems can estimate injury severity, identify incidents with serious-injury or fatality potential, generate task-specific guidance and combine written information with images, sensors and equipment data.

The review also identifies important limits. Models depend on the quality and coverage of historical records and may produce biased or fabricated outputs. The authors recommend using artificial intelligence as decision support, with occupational safety professionals retaining responsibility for safety-critical decisions.

Aim and Context

Workplaces generate large quantities of written safety information, including investigation narratives, inspection comments, hazard observations, exposure records and job descriptions.

Much of this information is unstructured, making it difficult to compare across large datasets.

The review examined how artificial intelligence can extract useful information from these records, how applications have developed from text mining to generative models, and what controls are needed when these systems influence occupational safety decisions.

Methodology

The authors conducted a structured review, described in the paper as a scoping review.

The Web of Science Core Collection was the main database, supported by Google Scholar searches and citation tracking. Studies published from 2013 to October 2025 were considered.

A broad search produced 4,897 records. After screening and full-text assessment, 123 primary studies were included.

The studies covered conventional machine learning, deep learning, natural language processing, transformers, large language models, retrieval-augmented generation, computer vision and vision-language systems.

A narrative synthesis was used because the technologies, datasets and performance measures differed too much for a meaningful meta-analysis.

The review did not report a formal risk-of-bias assessment or standardised evidence-grading process.

Safety Records Can Be Analysed at Scale

Accident and near-miss narratives often contain information about the task, equipment, environment, supervision and sequence of events that may be lost when an incident is reduced to a fixed code.

Natural language processing can analyse thousands of these reports and identify recurring patterns.

Automated systems can also reduce manual coding and improve consistency. They cannot correct poor source information. Incomplete descriptions, inconsistent terminology and under-reporting remain problems regardless of the model used.

Models Are Moving From Classification to Prediction

Earlier systems mainly classified accident types, extracted keywords and grouped similar incidents.

Newer deep-learning systems have been used to estimate injury severity and identify reports with serious-injury or fatality potential.

This may help organisations decide which incidents require closer investigation, including events where the actual injury was minor but the potential consequence was severe.

Performance needs careful interpretation. Severe events are often rare, so high overall accuracy can conceal poor results in the category that matters most. False negatives and performance on rare events need separate assessment.

Large Language Models Can Generate Safety Guidance

Large language models can summarise reports, answer questions, draft investigation material and produce guidance linked to particular tasks or hazards.

Retrieval-augmented generation can connect a model to selected accident records, standards or internal procedures. This may reduce unsupported answers and make the source of recommendations easier to trace.

The quality of the output still depends on the documents supplied. Outdated, incomplete or incorrect material can produce unreliable guidance.

Controlled document repositories, version management and professional review are therefore necessary where generated advice could affect safety decisions.

Artificial Intelligence Can Support Investigations

Aviation studies show how language models can extract human, technical, environmental and organisational factors from detailed incident narratives.

Similar approaches can help investigators search previous cases, identify possible contributing factors and organise information against recognised investigation frameworks.

The review does not support replacing the investigator. Interviews, site conditions and factual evidence still need to be checked by competent people. Generated explanations can appear convincing even when the evidence does not support them.

Applications Extend Across High-Risk Sectors

Construction studies include accident classification, extraction of risk factors, safety knowledge bases and task-specific guidance.

Chemical, process and mining applications analyse fires, explosions, confined-space incidents, equipment failures and accident pathways. Some combine text analysis with fault trees, Bayesian networks and other established safety methods.

Transport applications process road, rail, maritime and public-transport records. Multilingual capability is particularly relevant, although models trained mainly on English may perform less reliably with specialist terminology in other languages.

Artificial intelligence has also been used to classify occupations for health surveillance and to support personal protective equipment monitoring through vision-language systems.

Multimodal Systems Can Combine Several Sources

Some systems combine written reports with sensor readings, process information, machine status, images or video.

This could allow risk information to be updated as conditions change rather than waiting for a scheduled review.

These systems also create practical difficulties. Data sources need to be synchronised, false alarms managed and alerts connected to existing safety processes.

Camera-based systems raise further questions about privacy, surveillance, data retention and worker consultation.

Data Quality and Validation Remain Major Constraints

Artificial intelligence learns from existing records. If subcontractor incidents, minor events or concerns raised by poorly represented workers are missing, the model may reproduce those gaps.

Many systems have also been tested only on data from the organisation or sector used for training.

Good performance in one dataset does not establish that a model will work with different terminology, equipment, working practices or regulatory requirements.

Models should therefore be tested across different sites, time periods and workforce groups.

Hallucinations Create a Safety Risk

Large language models can generate plausible but incorrect information.

In occupational safety, this could mean incorrect control measures, invented legal requirements or failure to identify an important hazard.

Retrieval-augmented systems may reduce this risk but cannot remove it completely.

Generated findings should therefore be checked before they influence risk assessment, investigation or safety-critical decisions.

Takeaways for Practice

  • Organisations should start with a defined safety problem rather than adopting artificial intelligence without a clear purpose.
  • Incident reporting quality should be improved before historical records are used for training or evaluation. Model performance should also be tested on rare but severe events, not judged only by overall accuracy.
  • Where retrieval systems are used, the underlying documents should be controlled and kept current.
  • Artificial intelligence can help safety professionals process larger amounts of information, identify patterns and prioritise attention. Responsibility for interpreting the evidence and making safety decisions should remain with competent people.

Read the full research study here:

New Trends in the Use of Artificial Intelligence and Natural Language Processing for Occupational Risks Prevention. Natalia Orviz-Martínez , Efrén Pérez-Santín and José Ignacio López-Sánchez https://www.researchgate.net/publication/399585378_New_Trends_in_the_Use_of_Artificial_Intelligence_and_Natural_Language_Processing_for_Occupational_Risks_Prevention

Tags: research, occupational health, core health & safety, compliance