Deep DivesFebruary 6, 2026

Most AI Research Tools Are Just ChatGPT With Data

There's a wave of AI tools targeting equity research. Most share the same architecture: take an LLM, connect it to financial data, let analysts ask questions. It's a compelling demo. But it's not research infrastructure, and the distinction matters more than most people realize.

The Iceberg of AI Research Tools — most tools show only the surface while the real infrastructure lies beneath
Shams Hasan Rizvi
Shams Hasan Rizvi
KnowYourCompany.ai10 min read

TL;DR

Most AI equity research tools share the same flawed architecture: an LLM bolted onto financial data with no orchestration, workflow, or materiality layer in between. This creates a trust problem where analysts must manually verify every AI-generated answer, often spending more time than if they had done the research themselves. The competitive edge will belong to systems that compound intelligence through layered infrastructure, not to tools that merely answer questions faster.

There's a wave of new AI tools targeting equity research.

I've seen a lot of products, from startups and incumbents alike, that describe essentially the same architecture: take a large language model, connect it to financial data sources, and let analysts ask questions in natural language.

It's a compelling demo. Type a question, get an answer with citations. It feels like the future.

But it's not research infrastructure. And the distinction matters more than most people realize.

The Chatbot Trap

Here's the pattern I keep seeing: a product that lets you ask "What was Infosys's gross margin trend over the last 8 quarters?" and returns a nicely formatted answer with charts and source links.

That's useful. I won't pretend it isn't. It's faster than pulling up exchange filings manually or navigating a legacy terminal.

But ask a harder question: "Given Zomato's margin trajectory, commentary on delivery unit economics from the last two earnings calls, and the competitive pricing moves from Swiggy this quarter, is my bull thesis on contribution margin expansion still intact?" And these tools fall apart.

Why? Because answering that question requires more than retrieval. It requires orchestration: the ability to gather data from multiple sources, assess materiality, reconcile conflicting signals, evaluate what's changed since the thesis was formed, and synthesize it into a judgment-ready summary.

Most AI tools can do the first step. Almost none can do the rest.

Let me make this concrete.

What "Falling Apart" Actually Looks Like

I watched an analyst test three different AI research tools with the same prompt last quarter. The prompt was straightforward, something like: "Has anything in the last 30 days changed the bull case for HDFC's net interest margin expansion?"

Here's what happened:

Tool A pulled the most recent earnings transcript and summarized the margin-related commentary. It cited the CFO's prepared remarks about "continued operational efficiency" and returned a confident answer: the bull case is intact. What it missed: a key NBFC partner had disclosed rising NPAs in their own BSE filing two weeks later. That disclosure hadn't made it into any earnings commentary yet, but it directly threatened the asset quality story underpinning the NIM expansion thesis.

Tool B was more comprehensive. It pulled the earnings transcript, the quarterly filing, and three sell-side notes. But it treated all the data with equal weight. The NBFC partner's NPA disclosure appeared as one bullet point among twelve, sandwiched between a management quote about "strong demand visibility" and a consensus estimate revision from a broker. No materiality assessment. No flag that this single data point potentially invalidated the core thesis.

Tool C returned a well-structured answer that looked impressive on first read. But when the analyst checked the sourcing, one of the key citations was from a filing that was two quarters old. The tool had retrieved relevant-looking content without verifying temporal context. In equity research, a data point from Q2 presented as current evidence is worse than no data at all; it creates false confidence.

Three tools. Three different failure modes. But the same underlying problem: none of them had a mechanism for evaluating whether their answer was actually reliable.

The Missing Layers

I think about AI research systems as a four-layer pyramid:

Layer 1: Data Platform. The foundation. Structured and unstructured financial data, ingested, normalized, and made queryable. Earnings transcripts, filings, estimates, market data, news, alternative data. This is table stakes, and frankly, several companies do this well.

Layer 2: Orchestration. This is where most tools have a gaping hole. Orchestration is the intelligence layer that determines what's material, what's changed, what conflicts with existing analysis, and what deserves attention. It's the difference between "here are 50 data points" and "here are the 3 that matter for your thesis right now."

In the example above, an orchestration layer would have flagged the supplier price increase as a high-materiality event, because it directly intersects with the margin expansion thesis the analyst is tracking. It wouldn't have buried it in a list. It would have led with it.

Layer 3: Workflows. Persistent analytical processes that compound over time. Not one-off Q&A, but continuous monitoring, thesis tracking, comparative analysis across a coverage universe, and pattern detection across multiple companies and quarters.

A workflow layer would have known that the analyst has been tracking the margin expansion thesis for three quarters. It would have automatically connected the supplier filing to the existing thesis context, flagged the potential conflict, and presented it with the historical margin trajectory already assembled.

Layer 4: Agents. AI that can execute complex research tasks, but only when grounded in verified data (Layer 1), filtered through materiality assessment (Layer 2), and operating within a structured analytical framework (Layer 3).

Most AI tools on the market today are Layer 4 bolted directly onto Layer 1. An agent connected to a data source. It's like building a roof directly on the ground; it'll keep the rain off for a while, but it's not a building.

The Trust Problem

Here's why this architecture gap isn't just a technical quibble. It's a trust problem.

When an analyst asks a question and an AI returns an answer, the analyst needs to trust that answer enough to put capital behind it. Not "this sounds plausible" trust. "I'm comfortable presenting this to my PM and recommending we adjust our position" trust.

That's a high bar. And it should be.

An AI that retrieves data and generates a response has no mechanism for evaluating whether that response is reliable. It doesn't know if the data is stale. It doesn't know if there's a conflicting signal it missed. It doesn't know if the context has changed since the thesis was formed. It just generates the most plausible-sounding answer.

I've started asking analysts a simple question during our user research: "When an AI tool gives you an answer, what do you do next?"

The answer is almost always some version of the same thing: "I go verify it myself."

Think about what that means. The tool that was supposed to save the analyst time has actually added a step to their workflow. They ask the AI, get an answer, and then spend 20-30 minutes checking whether the answer is trustworthy. In some cases, it would have been faster to just do the research manually.

This is the paradox of low-trust AI: it creates the illusion of productivity while actually imposing a verification tax.

The way out isn't a better language model. It's a better architecture. An agent without a data platform is just a fast talker. An agent with a data platform, orchestration, and workflows is a research system that compounds intelligence and earns trust through transparency, not confidence.

What Trust Actually Requires

When I talk to analysts who have developed genuine trust in a research system, human or AI, they describe the same set of properties:

Provenance. Every claim traces back to a specific source, with a timestamp. Not "according to the earnings call" but "Q3 2025 earnings call, CFO prepared remarks, page 4, paragraph 3." If the analyst can't verify the source in under 30 seconds, the system hasn't earned trust.

Materiality awareness. The system distinguishes between noise and signal, not by volume, but by relevance to the analyst's specific context. A 2% revenue beat is noise for one company and a thesis-changing event for another. The system needs to know the difference.

Temporal integrity. Data is presented with clear temporal context, and the system never mixes current and stale information without explicit flagging. The Tool C failure mode I described earlier, presenting old data as current evidence, is one of the fastest ways to destroy trust.

Conflict surfacing. When data points contradict each other, the system flags the conflict explicitly rather than resolving it silently. A language model's instinct is to synthesize conflicting information into a coherent narrative. In equity research, the conflict is the insight. Smoothing it over is a disservice.

Sufficiency honesty. When the system doesn't have enough context to form a reliable view, it says so. This is perhaps the hardest property to build, and the one most AI tools skip entirely. In a domain where being wrong has real financial consequences, "I don't have enough evidence" is more valuable than a confident guess.

These aren't nice-to-haves. They're the minimum requirements for a system that an analyst can actually rely on.

Being Generous and Being Honest

I want to be clear: the current wave of AI tools is a genuine step forward from the status quo. Analysts have been drowning in manual processes for decades, copying data between terminals, reformatting spreadsheets, searching through hundreds of pages of filings for a single data point. Anything that reduces the time from "I have a question" to "I have an answer" is valuable.

And the teams building these tools are talented. The engineering required to parse financial documents, build reliable data pipelines, and generate coherent natural language responses from complex datasets is genuinely impressive work.

But I think the industry needs to think bigger.

The question isn't "Can AI answer my question?" It's "Can I trust this answer enough to put capital behind it?"

And the answer to that second question requires infrastructure that most tools haven't built yet. It requires orchestration that separates signal from noise. It requires workflows that maintain analytical context over time. It requires agents that know when to say "I don't have enough evidence" instead of generating a confident-sounding guess.

Where This Separates

The competitive landscape in AI for equity research is going to separate quickly into two tiers: tools that answer questions, and systems that compound intelligence.

Both are useful. But only one will become essential infrastructure.

Tools that answer questions will compete on speed, data breadth, and interface design. They'll get faster and more comprehensive, and they'll save analysts real time on routine information retrieval. That's a valuable market.

But systems that compound intelligence, those that maintain analytical context over quarters, proactively surface material changes, and build trust through architectural integrity, will become the operating layer that firms can't turn off. They'll be the difference between an analyst who's fast and an analyst who's reliably right.

The firms and analysts who recognize this distinction early will have a meaningful edge. Not because they have better AI, but because they have a better system for turning AI into decisions.

And in equity research, the ability to turn information into trusted decisions faster than anyone else is the only edge that compounds.


Other posts in the Deep Dives series:

Enjoyed this article?

Get more insights like this delivered to your inbox weekly.

Continue Reading

Logo

Unlock financial AI for your firm