AI in Requirements Engineering at BMW

Raiqon AI is now serving thousands of engineers at BMW, demonstrating large scale enterprise deployment with millions of items across the V-Model.

AI in Requirements Engineering at BMW

BMW operates AI-assisted requirements engineering at a scale that changes the nature of the problem. Around 5.5 million requirements artifacts exist inside the engineering landscape, and the dataset grows by roughly 120,000 artifacts every month. Around 7% of all BMW employees use Codebeamer.

Requirement specifications per year
64%

Percentage of items in Codebeamer that are requirements

Data growth per month

This environment exposes the limits of generic AI tools quickly. Copying artifacts into chat interfaces and manually transferring results back into engineering systems does not work at this scale because the overhead grows continuously with the number of users and artifacts. Manual workflows introduce friction into engineering processes, weaken traceability, complicate permission handling, and increase operational cost.

The central question therefore changes. The issue is no longer whether large language models can generate useful engineering text. The issue is how AI systems operate inside large engineering organizations with existing toolchains, compliance obligations, and millions of interconnected artifacts.

Download the Presentation Slides from 2026 ProSTEP Symposium

Enterprise AI Has Different Constraints

Public AI discussions still focus heavily on conversational interfaces. Enterprise engineering environments have different priorities.

BMW’s engineering landscape includes safety and compliance requirements, established ALM workflows, verification and validation activities, large engineering organizations, and continuous data growth. These constraints shift the focus away from prompting and toward integration, infrastructure, and operational behavior.

An engineer working on isolated requirements inside a pilot project can tolerate manual interaction with an external AI tool. Large organizations cannot. Every additional export step, copy operation, or manual review activity multiplies across thousands of users and millions of artifacts.

The economics also change at this scale. Small inefficiencies become operational problems once they affect entire engineering organizations. Enterprise concerns like GPU consumption, inference cost, permission handling, or throughput become important concerns.

AI deployment in engineering therefore resembles enterprise infrastructure deployment far more than consumer software adoption.

Scale Changes the Economics

Most AI demonstrations operate on relatively small datasets. BMW does not. The operational figures illustrate the difference clearly:

For a subset of 2 million requirements, the platform processed:

  • 1.8 billion input tokens
  • 5.9 billion output tokens
  • Formal analysis of 2 million requirements
  • 4,000 GPU hours completed within 48 hours

At this scale, inference efficiency becomes a deployment requirement rather than a technical detail.

A workflow that appears affordable during a pilot project can become prohibitively expensive once expanded to millions of artifacts and thousands of users. Large engineering organizations therefore evaluate AI systems differently from small teams experimenting with public chat tools.

Throughput also becomes critical. Large engineering environments cannot wait days or weeks for large-scale analyses because engineering workflows continue moving while the system processes data.

The deployment also emphasized efficiency improvements compared to alternative approaches. The precise benchmark matters less than the broader conclusion. Enterprise AI deployment becomes an infrastructure economics problem very quickly once organizations move beyond isolated pilots.

Integration Matters More Than Chat

BMW integrated AI capabilities directly into Codebeamer workflows instead of treating AI as a separate assistant system.

This changes how engineers interact with the system. Context no longer has to be assembled manually because the AI system already knows the artifact structure, relationships, permissions, and engineering context.

The deployment followed three stages: Analyze, Propose, and Execute.

Analyze covers activities such as weak-word detection and formal requirement analysis. Propose generates concrete suggestions, alternative formulations, or candidate trace links. Execute performs approved actions directly inside the engineering environment.

The business impact increases sharply across these stages because the AI system moves from isolated analysis toward direct workflow participation. Standalone analysis tools save local effort, while integrated execution changes the engineering process itself.

This becomes visible in activities such as traceability management. Maintaining links across requirements, architecture, implementation, and verification artifacts consumes enormous engineering effort when performed manually. AI-assisted traceability changes the economics of this activity because links can be analyzed and proposed across millions of artifacts inside the existing workflow environment.

Determinism Became a Central Requirement

One topic receives surprisingly little attention in public AI discussions despite becoming central in enterprise engineering environments: determinism.

Most generative AI systems are non-deterministic. The same input can produce different outputs across executions. This behavior works for brainstorming and creative interaction, but engineering environments operate under different conditions.

Engineers lose trust in systems that behave inconsistently. Validation becomes difficult when outputs fluctuate between executions. Benchmarking becomes unreliable because organizations can no longer compare results consistently across teams, projects, or optimization cycles. Compliance discussions also become harder because engineering organizations must explain how results were generated.

Raiqon addressed this through deterministic execution, where identical input produces identical output.

This changes how AI systems behave operationally. Validation becomes repeatable, measurements become stable, and optimization becomes measurable because organizations can reproduce analyses reliably across executions and teams.

Public AI discussions rarely focus on determinism because conversational systems tolerate variation naturally. Enterprise engineering environments do not.

Deterministic behavior enables repeatable validation, measurable optimization, reproducible analyses, and higher organizational trust. For engineering organizations, that difference is substantial.

Engineering Data Becomes Usable Again

Large engineering organizations possess decades of engineering knowledge stored across requirements databases, specifications, test reports, source repositories, and verification artifacts. Much of this information remains operationally inaccessible because the datasets are too large, fragmented, and inconsistent for manual analysis.

AI changes this situation only if organizations can process sensitive engineering data inside secure environments.

BMW therefore emphasized deployment models that operate within enterprise boundaries instead of relying entirely on external cloud services. Domain-specific and smaller AI models make this practical because they can be fine-tuned and operated on-premise with strictly confidential engineering IP.

Integrated AI systems can analyze these datasets directly inside the engineering environment. Historical engineering data becomes searchable, linkable, and operationally useful again.

The value of enterprise AI therefore does not lie only in generating new content. It also lies in making decades of existing engineering knowledge accessible across large organizations.

AI Expands Beyond Requirements Authoring

Many AI initiatives in requirements engineering focus narrowly on writing support. The emphasis lies on generating or reformulating requirements, activities on the left side of the V-Model.

Large engineering organizations spend enormous effort elsewhere. Verification, validation, consistency checking, traceability maintenance, and evidence generation dominate engineering effort in many regulated environments.

Left side of V-Modell

  • Writing requirements
  • Improving wording
  • Generating specifications
  • Structuring information

Right side of V-Modell

  • Validation
  • Traceability
  • Discovery
  • Verification support
  • Cross-artifact consistency analysis

AI systems integrated directly into engineering workflows can support these activities at scale. Traceability analysis across millions of artifacts becomes feasible. Cross-artifact consistency checks become feasible. Large-scale formal analysis becomes feasible.

This changes the role of AI inside engineering organizations. The system no longer acts primarily as a text generator. It becomes part of the engineering infrastructure itself.

The bottleneck is maintaining alignment, consistency, traceability, and verification across extremely the V-Model in large engineering systems.

Enterprise AI Deployment Looks Different

BMW’s deployment illustrates a broader pattern emerging across enterprise engineering.

The central challenge is no longer text generation quality. Large language models already generate useful engineering text. The harder questions lie elsewhere.

Organizations must determine whether AI systems integrate into existing engineering environments, operate under enterprise permission models, scale economically, support compliance processes, and behave predictably under operational conditions.

These requirements determine whether AI remains a pilot project or becomes part of engineering infrastructure.

Public perception still revolves heavily around chat interfaces. Large engineering organizations increasingly evaluate different properties instead: deterministic behavior, integration depth, throughput, inference economics, and support for verification workflows.

That shift changes how AI systems for engineering environments are designed and evaluated.

Related Posts