Skip to content
SevenUpdatesNews ยท 24/7
Crypto

Coinbase AI Models Miss More Fraud

Coinbase's AI model upgrades caught fewer fraudulent payments, challenging assumptions about model improvements.

๐Ÿ’ฌ ๐• f in
A graph showing the decline in recall rates for Coinbase's AI models
A graph showing the decline in recall rates for Coinbase's AI models

Key Takeaways

  • Coinbase's AI model upgrades caught fewer fraudulent payments and had lower recall rates
  • The company's findings highlight the importance of thorough testing and evaluation of AI models in payment screening
  • Payment providers should test the configuration first, then evaluate changed prompts or thresholds separately, with latency, reliability and cost alongside detection quality

Coinbase's AI Model Upgrades Miss More Fraud

Coinbase reported that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service. The findings challenge the assumption that upgrading a model improves an existing payment screener.

The company's evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. Each candidate reviewed recent transaction behavior under fixed guidance and the same policy for turning risk classifications into decisions.

Background and Context

Crypto fraud detection is a critical aspect of the digital currency ecosystem, with millions of transactions taking place every day. As the use of cryptocurrencies continues to grow, the need for effective fraud detection systems becomes increasingly important. Coinbase, one of the leading cryptocurrency exchanges, has been at the forefront of developing and implementing AI-powered fraud detection models.

Key Findings

Coinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall.

  • Sonnet's recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points.
  • Opus's recall declined 0.8 points.
  • GPT showed improved precision, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points.

The replay does not establish customer losses from deploying those versions. Coinbase also said it could identify the regressions without establishing their cause.

Expert Perspective

According to experts in the field, the findings highlight the importance of thorough testing and evaluation of AI models in payment screening. The fact that newer models caught fewer fraudulent payments and had lower recall rates suggests that the models may not be effective in detecting certain types of fraud.

Implications for Readers in India

The implications of Coinbase's findings are significant for readers in India, where the use of cryptocurrencies is growing rapidly. As the Indian government continues to develop regulations for the cryptocurrency market, the need for effective fraud detection systems becomes increasingly important. Readers in India should be aware of the potential risks associated with cryptocurrency transactions and take steps to protect themselves from fraud.

Limitations and Future Directions

The SR-Fraud researchers say the proprietary dataset cannot be released, restricting independent replication and generalization. In its Oct. 8 disclosure, Coinbase reported that a post-trained Qwen3.5-9B model exceeded Opus 4.5 across four fraud-detection metrics.

Separately, production measurements put median end-to-end LLM-request latency at 0.683 seconds versus 1.515 seconds for Opus 4.5, a 55% relative reduction. Faster inference and stronger benchmark detection came from different evaluations.

Recommendations for Payment Providers

Coinbase recommends testing the configuration first, then evaluating changed prompts or thresholds separately, with latency, reliability and cost alongside detection quality.

For payment providers, the upgrade question is whether a candidate improves fraud coverage under their actual decision setup. The company's findings highlight the importance of thorough testing and evaluation of AI models in payment screening.

What to Watch Next

As the cryptocurrency market continues to evolve, readers should watch for further developments in AI-powered fraud detection systems. The use of machine learning and deep learning algorithms is likely to become more prevalent, and the development of more effective fraud detection models will be critical to the growth and adoption of cryptocurrencies.

Timeline of Events

Coinbase's historical replay found lower fraud-case and fraud-value detection across three model upgrades under a fixed decision policy. The company's evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions.

  • Oct. 7: Coinbase reported that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service.
  • Oct. 8: Coinbase disclosed that a post-trained Qwen3.5-9B model exceeded Opus 4.5 across four fraud-detection metrics.

Conclusion

In conclusion, Coinbase's findings highlight the importance of thorough testing and evaluation of AI models in payment screening. The company's recommendations for payment providers emphasize the need to test the configuration first, then evaluate changed prompts or thresholds separately, with latency, reliability and cost alongside detection quality.

Frequently Asked Questions

What were the key findings of Coinbase's historical replay?

The replay found lower fraud-case and fraud-value detection across three model upgrades under a fixed decision policy.

What are the implications of Coinbase's findings for readers in India?

The implications are significant, as the use of cryptocurrencies is growing rapidly in India and the need for effective fraud detection systems becomes increasingly important.

What are Coinbase's recommendations for payment providers?

The company recommends testing the configuration first, then evaluating changed prompts or thresholds separately, with latency, reliability and cost alongside detection quality.

Share this story WhatsApp X Facebook LinkedIn
24
24SevenUpdates Editorial Desk

Our newsroom tracks India and the world around the clock, turning verified reporting into clear, fast explainers. Read how we source, verify and correct our stories in the editorial policy, or contact the desk about this story.