Press Release Checkmarx Fusion: Hybrid Scanning Delivers the Most Complete Vulnerability Detection Available Read Now
Gartner® Checkmarx Named a Leader in the 2026 Gartner® Magic Quadrant™ for Software Supply Chain Security Get the Report
Outlook Report The Future of Application Security in the Era of AI Download Now
Webinar The AppSec Bottleneck Has Moved Downstream: Why visibility is no longer enough in the age of AI-generated code Watch Now
Latest Innovations
Checkmarx for Developers
Partners
Blog
Research

Risk, Return, and Remediation: An Investment Strategy for Application Security

Learn why a hybrid application security strategy combining deterministic tools and AI delivers better protection, lower costs, and faster remediation for AI-generated code.

Don’t Put All Your Code in One Basket.

Every seasoned investor knows to never put your money behind a single asset.   
A successful investment strategy is diverse.  
 
No one puts all their money in stocks: markets correct. Real estate goes through cycles of boom and bust. Cash may feel safe, but inflation erodes its value over time. Having different investment assets helps offset volatility, so your portfolio as a whole stays resilient, even if that meme coin you bought on a whim takes a nosedive.  
 
The same principle of diversification applies to our work in application security. Staying ahead of emerging threats and bad actors in the era of AI-generated code demands more than a single line of defense. It requires a layered, deterministic, and probabilistic approach. Each approach has limits on its own. Together, they work like a diversified portfolio. 

The Growing Gap Between Functional and Secure 

This is the essence of our Checkmarx strategy, and we’re thrilled to see our vision backed by independent, hard data. We sponsored recent research published in The Weather Report, a non-profit AI security and safety lab led by Dr. Ilya Kabanov, former founder of AI Safety & Security at Google and a research affiliate at MIT Sloan Cybersecurity. While Checkmarx funded the study, Kabanov and his team retained full control over design, methodology, evaluation, and conclusions.

In his research, Kabanov re-ran two established security benchmarks: a snippet-level test, and a broader agentic test built around real open-source repositories. He tested both against frontier models in their native CLIs: Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, and Gemini 3.5 Flash.

The results? These models now produce working code 83–95% of the time, but secure- 
and-functional code, or output that both works and passes security tests, still reaches only 24–36%.

That finding alone should give any AppSec manager pause. But Kabanov’s research went even further. Adding inference-time security steps, or threat modeling before coding, security review after, raises the secure-and-functional rate from a baseline of 24-36% to 47-56%.

But this improvement comes at a steep cost. Running a task through the full intervention ladder uses roughly four to five times the tokens required to write the code alone, while still leaving roughly half of all tasks insecure.

That’s like paying a premium for a guess dressed up as a guarantee. But even then, Kabanov isn’t convinced that better code generation leads to more secure code, noting that “functional capability is a weak predictor of security.” In his findings, one of the top-ranked coders actually finished last on security

LLMs Alone Can’t Keep the Codebase Secure

These are the reasons we built the Checkmarx One unified platform as the most efficient, cost-effective way to deliver security at the speed of frontier LLMs. It’s a deterministic floor for ground truth and auditability, offering AI-augmented reasoning for context and novel patterns, with security context (reachability and business risk) layered on top.

Because while the leading LLMs are exceptionally good at code generation, when it comes to protecting that code, an “LLM-only” approach isn’t enough. The smart investor never bets on a single asset. The smart AppSec team doesn’t bet on a single model. 
 
Our deterministic tools apply consistent rules and proven detection logic, producing repeatable results that teams can trust and verify. That’s the core of our Checkmarx vision: a hybrid platform that pairs certainty (deterministic ground truth) with possibility (probabilistic AI reasoning). 

Relying on the same model to check the code that it writes is an inherent conflict of interest. The Chief Financial Officer doesn’t audit her company’s financials. We champion independent verification in AppSec: a system of checks and balances to keep any single point of failure from compromising the stack.  

Tokenization Savings Is Only the Start

Not only are there cost savings to be realized by leveraging a hybrid AppSec strategy, but when issues are addressed earlier in the development lifecycle, the cost to fix them goes down. Way down.  
 
In a recent talk I gave, “When Code Secures Itself: The Rise of Agentic AI in Application Security,” I outlined four control points across the AI development lifecycle and a simple, uncomfortable truth: the cost of fixing a vulnerability rises roughly 10x at every stage you wait. 
 
Fixing a vulnerability in the IDE is roughly 10x cheaper than in the build pipeline, 100x cheaper than in the AI supply chain, and 1,000x cheaper than in runtime.  
 
Yet only one in five developers embed security at the point of code creation, according to the 2,350 AppSec professionals we surveyed for our Future of Application Security Report. 
 
More than 80% of AppSec is conducted at defined stages after the code already exists, or worse, reactively once incidents surface. And the later a flaw is found, the more it costs in time, money, and exposure. 

Show Me the Metrics

An F1 score measures a scanning engine’s fidelity, how well it finds real vulnerabilities without burying teams in false positives.

In head-to-head testing across seven real production codebases, Checkmarx One’s hybrid engine achieved an F1 score of 0.64 – more than three times the industry average of 0.20 – across competing approaches that Checkmarx evaluated.

We are building on this momentum by helping AppSec teams address the new challenge of AI-generated code: shifting the focus from discovery and “mass-detection” to “mass remediation” to fix vulnerabilities quickly and consistently, whether newly-created or languishing in the backlog. 
 
Our newly-launched agentic experiences – Triage Assist, Remediation Assist, and Developer Assist – are purpose-built for this moment.

Triage Assist cuts through vulnerability noise using reachability, exploitability, policy context, and application risk to surface what deserves attention first. Remediation Assist then generates contextual, merge-ready fix recommendations that fit how developers already work. Developer Assist 2.0 is a fully autonomous virtual assistant working alongside developers in the IDE, detecting issues, generating fixes, and iterating until code is clean.

The result: fewer hours lost to manual analysis, fewer handoffs between AppSec and development, and faster remediation across the ADLC. This drives better credibility for AppSec teams, and more time for developers to do what they do best: build. 
 
But knowing how to find and fix vulnerabilities faster only matters if you know what you’re securing in the first place. That’s the gap our newly-launched Checkmarx AI Inventory closes.  
 
AI Inventory gives teams continuous, deterministic visibility into every AI component running in their code (like models, agents, MCP servers, SDKs,) and generates an AI-BOM (AI Bill of Materials.) This provides policy controls and audit-ready documentation for every AI component discovered, helping teams manage assets, monitor risk, and stay ahead of emerging compliance requirements.

The Portfolio Approach to Application Security

The best portfolios don’t win by picking one great asset. They win by pairing complementary strengths to deliver returns no single position could achieve alone.

In application security, that means combining the fluency of frontier AI with the certainty of deterministic enforcement.

Working in balance. Proven by the numbers. Built to scale with whatever comes next. 

That’s the Checkmarx strategy. Built for today’s threats. Ready for whatever comes next. Get a personalized AppSec demo from Checkmarx today. 

Learn More About What’s New from Checkmarx

Checkmarx Triage Assist  
Checkmarx Remediation Assist  
Checkmarx Developer Assist 
Checkmarx AI Inventory 
 
Download the 2026 Checkmarx Future of Application Security Report

Tags:

Agentic AI

Application Security Testing

AppSec