Back to blog
AI News

Third party cybersecurity evaluations and what UK businesses should do next

OpenAI is adding safeguards after third party cyber evaluation incidents. UK businesses using AI should tighten testing controls, expand threat awareness, and align model use with clear risk processes.

What changed in third party cybersecurity evaluations

OpenAI says there have been recent issues in third party cybersecurity evaluation work involving its models, and it has moved to strengthen how models are tested and evaluated. The focus is on adding safeguards so evaluation activities are more reliable and risk aware, rather than treating testing as a purely technical step.

Why this matters for businesses

For UK organisations that rely on AI in operational workflows, evaluation incidents are a signal to treat model testing and monitoring as a continuing control, not a one off checkbox. Even if your use case is internal, you still need to know how model behaviour can shift when conditions change, and how evaluation settings might surface different risks.

Practical next steps for AI teams in operations and customer workflows

  • Review your current AI evaluation process, including who runs tests, what gets tested, and what success criteria mean for safety and compliance.
  • Add explicit controls around model access during testing, so experiments do not accidentally expand capabilities beyond what your business intends to allow.
  • Update your risk assessment to cover third party evaluation scenarios, including how results should be interpreted and how you will respond if testing reveals new issues.
  • Strengthen ongoing monitoring for model outputs in your real workflows, especially in areas tied to security sensitive customer information or operational decision making.
  • Document what changed after incident learnings and ensure relevant teams, including product, operations, and governance, know which controls were added and why.

How New Era AI recommends approaching this in a UK context

We suggest starting with a short internal gap review of your AI controls, then tightening evaluation governance and output monitoring, and finally updating operational playbooks so teams know exactly how to react when tests uncover safety or security concerns. This keeps adoption practical while reducing avoidable risk.

Third party cybersecurity evaluations and what UK businesses should do next | New Era AI