The headline hit my feed yesterday: Bank of America launches an AI tracking tool covering model intelligence and costs. My first reaction wasn’t excitement—it was a chill. Because when a bank decides to become the umpire of AI performance, we need to ask: Who tracks the tracker?
Context: The Fragmented AI Evaluation Landscape
For the past two years, the AI model evaluation space has been a chaotic bazaar. Independent platforms like LMArena, Artificial Analysis, and Hugging Face’s Open LLM Leaderboard provide raw benchmarks. Developers cross-reference API pricing from Vellum. But there’s no single, trusted source—especially for institutional investors. That’s the gap Bank of America (BoA) is trying to fill.
Their tool reportedly aggregates two key dimensions: model intelligence (likely based on public benchmarks like MMLU, HumanEval, MATH) and cost (per-token API pricing). The output is a comparative score that helps enterprises decide which model to use. On the surface, this sounds useful. But as a crypto evangelist who’s spent years arguing that financial gatekeepers are the problem, not the solution, I see a darker story.
Core: The Centralization Trap
Let’s cut through the hype. This tool is not a technical breakthrough—it’s a power play disguised as a service. BoA’s core advantage is its institutional client network. They can push this tracker directly into the hands of CTOs, CFOs, and investment committees. The problem? The tracker’s methodology is opaque. We don’t know how intelligence is weighted, what benchmarks are used, or how costs are defined (API only or total cost of ownership? Training costs?).
Based on my experience building a crypto education platform and analyzing DeFi protocols, I’ve learned one thing: centralized ratings always favor the raters’ interests. During the 2017 ICO frenzy, I saw how “independent” ratings from banks and research firms were often influenced by which projects were paying for their services. This tool is no different. BoA is a major investment bank for AI companies. If they give a client’s model a low score, they risk losing that client. If they boost a non-client, they might win future business. The conflict of interest is baked in.

The real innovation would be a decentralized, on-chain model evaluation registry—where benchmarks are verified by smart contracts, and scores are determined by community consensus, not a bank’s internal committee. We have the technology: zero-knowledge proofs can verify model inferences without revealing proprietary data; decentralized oracle networks can aggregate multiple benchmark sources. But instead, we get a walled garden from a bank that wants to be the gatekeeper of AI trust.

Contrarian: The Pragmatic Defense
I can hear the rebuttal: “But David, this tool actually helps enterprises adopt AI faster. It reduces information asymmetry. It’s a net positive for the industry.” I don’t fully disagree. Standardized evaluation can accelerate ROI calculations and push model providers to compete on both price and performance. That’s good for the end user. But the devil is in the details.
Consider the risk of indicator simplification. Reducing “model intelligence” to a single score ignores crucial dimensions: safety, bias, robustness, and real-world performance. A model that scores high on MMLU might fail catastrophically in a medical context. BoA’s tracker won’t capture that nuance—it will just give a number. And enterprises will treat that number as gospel. We saw the same thing happen in crypto with centralized exchange ratings: they created a false sense of security, leading to billions in losses when the ratings were wrong.
Moreover, the tracker’s update frequency is a concern. AI models are evolving weekly. If BoA updates quarterly, the data becomes stale. The tool’s value decays rapidly. Yet, once it’s embedded in procurement workflows, it becomes sticky—hard to replace even when outdated.
Takeaway: Trust Is No Longer a Promise; It’s a Protocol
Bank of America’s AI tracker is a symptom of a larger problem: we’re handing over the keys to AI evaluation to the same institutions that failed us in 2008. The solution isn’t to ban the tool—it’s to build a better alternative.
Code is law, but empathy is the interface. We need a decentralized, transparent, and community-driven model evaluation standard. One where every benchmark is auditable, every score is derived from verifiable data, and every user can contribute. Until then, treat BoA’s tracker as what it is: a marketing tool for their investment banking business, not a neutral arbiter of AI truth.
I learned to stop preaching and start listening, but sometimes the signals are loud enough. The pivot wasn’t from hype to substance—it’s from centralized trust to decentralized verification. The future of AI evaluation belongs on-chain, not in a bank’s internal dashboard.