Why ChatGPT and Claude get your Amazon and Shopify numbers wrong

Paste an export into ChatGPT or Claude and the answer sounds right but is not. Ten ways AI misreads Amazon and Shopify data, and what fixes each one.

By Justin Maddahi · · 8 min read

The short answer

AI tools get Amazon and Shopify numbers wrong because the exports do not say what their numbers mean. The model cannot tell ordered sales from cash paid out, or a half-finished month from a closed one. It cannot tell one product listed five ways from five. So it guesses without saying so. The fix is data that labels its revenue basis, period and caveats, and flags sources that disagree.

ChatGPT and Claude get Amazon and Shopify numbers wrong because the exports do not say what their numbers mean. The files do not say which revenue they contain or which periods are still open. They do not say which rows are the same product, or which ad platform can see which checkout. So the model guesses, and it does not tell you it guessed.

Connecting a raw MCP server (a connector that lets an AI tool query a system directly) does not change this. It gives the model fresher rows, not the meaning of the rows.

Why a capable model still gets it wrong

A column called sales is a number without a definition. Your finance lead knows three things the file does not say. The Amazon figure is ordered sales before refunds. September is only half over. And “Lavender Hand Soap 12oz” and “Lavender hand soap - 12 oz” are one product. The model knows none of that unless the data says so. It does what a new hire does on the first morning. It takes every column at its name and adds things up.

The traps below are ones we have seen in real multi-channel brand data. For each one: what the AI says, why it is wrong, and what the data needs so the mistake cannot happen.

One question, five revenue numbers

Ask “what was our revenue in August?” with these exports in the chat. Illustrative numbers.

Source Figure What it actually is
Amazon Business Reports $120,000 Ordered product sales: the value of units ordered, before refunds
Amazon settlement report $81,000 Cash Amazon paid out after fees, refunds and other charges, for periods that do not match the calendar month
Shopify gross sales $60,000 Price times quantity, before discounts and returns
Shopify net sales $52,000 Gross sales minus discounts and returns
Shopify total sales $55,500 Net sales plus taxes, shipping and fees

The Shopify lines follow Shopify’s own sales report definitions.

A typical AI answer is “August revenue was $180,000.” That adds Amazon ordered sales to Shopify gross sales. Neither is net revenue, and the Amazon half is not money you received. A useful answer names its basis. It gives $120,000 of Amazon ordered sales and $52,000 of Shopify net sales. It reports Amazon cash separately. Every figure in the table is correct. Only a labeled one can be used.

Ten ways AI gets ecommerce numbers wrong

1. Gross, net and settlement revenue mixed together

What the AI says: “Revenue grew 22% month over month.”

Why it is wrong: Last month’s file was Amazon ordered sales. This month’s was a settlement export, Amazon’s statement of the cash it paid you. Or Shopify gross sales were in one file and net sales in the other. The growth is a change of definition. Settlement amounts also arrive on Amazon’s timetable. Amazon’s documentation notes that settlement reports cannot be requested or scheduled. They are generated automatically. So a settlement “month” rarely matches a calendar month.

What the data needs: every revenue figure stamped with its basis: orders, settlement, or store gross or net sales. Different bases are never joined into one trend line.

2. A month still in flight compared with closed months

What the AI says: “September sales are down 48% versus August.” Asked on September 15.

Why it is wrong: Half a month is being compared with a whole one. Even yesterday is not final. Amazon orders can sit in Pending status for days, and refunds land after the sale.

What the data needs: every period labeled closed or in flight. Trends use closed periods only, or compare the same days (September 1 to 14 against August 1 to 14).

3. Ad sales that keep changing after the week ends

What the AI says: “Last week’s Amazon ads returned 1.6x. Cut the budget.”

Why it is wrong: Amazon credits a sale to an ad if the purchase happens within a set time after the click. That time is called the attribution window. It is 7 days for Sponsored Products in Seller Central (Amazon’s seller dashboard). It is 14 days for Sponsored Brands and Sponsored Display (Amazon Ads attribution help). Amazon Attribution, the tool for measuring off-Amazon ads, also uses 14 days. Sales are credited back to the day of the click. So a week pulled the next morning keeps growing for up to two weeks.

What the data needs: ad figures labeled with their attribution window and whether it has closed, and recent weeks flagged as still counting.

4. Meta ads judged on checkouts the pixel cannot see

What the AI says: “Your Meta campaign has a 0.3x return. It is losing money.”

Why it is wrong: The Meta Pixel is tracking code on your own website. When an ad sends shoppers to an Amazon listing, they check out on Amazon. Your pixel is not installed there, so Ads Manager (Meta’s ad dashboard) never sees the sale. The campaign may be working. Meta simply cannot see where.

What the data needs: each campaign tagged with where it sends people: your store or Amazon. Pixel returns on Amazon-bound campaigns are labeled as not measuring the sale.

5. Customer counts built on Amazon buyer emails

What the AI says: “Only 4% of your customers buy on both Amazon and Shopify.”

Why it is wrong: Amazon does not give sellers a shopper’s real email. The address in your order data is an alias, a stand-in address, from Amazon’s Buyer-Seller Messaging service. It never matches the same person’s Shopify email, so the overlap reads near zero. In the data we see, the email can also arrive late or be missing on recent orders. So the latest weeks undercount customers.

What the data needs: Amazon customer counts labeled as alias-based, recent weeks marked as running low, and no cross-channel matching on email.

6. One product split into many rows

What the AI says: “Your best seller is the body lotion. The hand soap is fourth.”

Why it is wrong: The hand soap exists as one Shopify product and a refill bundle. On Amazon it is also a single bottle and a 3-pack, under different ASINs (Amazon’s product IDs). Each row is small. Together they are the top seller. The model ranked rows, not products.

What the data needs: a written map of which listings, variants and SKUs (your own product codes) belong to which product family. Apply it before anything is ranked.

7. Cohorts too young to judge repeat rate

A cohort is the group of customers whose first order fell in the same month.

What the AI says: “Repeat rate collapsed: 24% for January customers, 9% for July customers.”

Why it is wrong: By mid-September, January customers have had about eight months to reorder. July customers have had six to ten weeks. If a bottle lasts 90 days, no July buyer has run out yet.

What the data needs: every cohort carrying its age. Repeat rate is compared at the same age, such as 90-day repeat rate. Younger cohorts are marked too young to judge.

8. A partial first month that creates fake growth

What the AI says: “Shopify revenue is up 310% year over year in August.”

Why it is wrong: The store launched on August 19 last year, or the data connection only reaches back to that date. Last August holds 13 days of sales and this one holds 31.

What the data needs: a recorded first full month for every channel and source, and year-over-year comparisons that refuse a partial starting period.

9. Two tools disagree and the AI quietly picks one

What the AI says: “You had 1,240 orders in August.”

Why it is wrong: Shopify says 1,240. Google Analytics says 1,050, because browser tracking misses some shoppers. A subscription app counts renewals its own way. Given two numbers, a model will average them, take the first file, or take the one that fits its story. And it will say nothing.

What the data needs: a named source of record for each metric, meaning the one system you trust for that number. When two sources disagree by more than a set limit, a visible flag lets a person decide.

10. Missing COGS that flatters margin

COGS, cost of goods sold, is what the units you sold cost to make and bring into stock.

What the AI says: “Overall gross margin is 59%.”

Why it is wrong: Three of twenty products have no cost entered, and they are $18,000 of $100,000 revenue. The other products earn $82,000 on $41,000 of cost. Counting blank cost as zero gives 59%. If those three really cost $9,900, true margin is about 49%.

What the data needs: a cost per unit, with an effective date, for every product. Any margin figure that covers products with no cost says so. It also says how much revenue they represent.

How to check an AI answer this week

  1. Ask it to name the revenue basis of every figure. If it cannot, the file did not say.
  2. Ask which periods are complete. Compare any unfinished month on matching days only.
  3. Before cutting spend, pull Amazon ad results again once 14 days have passed since the last click.
  4. For campaigns that send traffic to Amazon, ignore pixel returns and look at Amazon sales.
  5. Give the model your product family list and have it group rows before ranking anything.
  6. Compare cohorts only at the same age.
  7. Check the first month of any year-over-year comparison for a partial start.
  8. Ask it to list every place two sources disagree instead of giving one figure.
  9. Ask which products have no COGS and what share of revenue they cover.

Doing this every time is tedious, and it lives in one person’s head.

Fix the data, not the prompt

A longer prompt fixes one conversation. The next chat, teammate or tool starts from zero. The lasting fix is data that carries its own meaning. Definitions are written once. Every number is labeled with its basis, period and caveats. Disagreements are flagged instead of settled in silence. That is what AI-ready data means, and the full guide walks through how to build it.

This is the problem Synthesis is built around. It joins a brand’s Amazon, Shopify, ads and subscription data into one warehouse (a single central database) per brand. It writes definitions down once, such as which listings form a product family and whether an ad sends people to Amazon or the store. It labels every number with its revenue basis, its period and caveats. Two examples are “this month is still in flight” and “Meta’s pixel cannot see Amazon checkouts”. When two figures for the same metric and period differ by more than 5%, the answer is flagged rather than one being picked. Claude and other AI tools reach it through MCP, locked to the signed-in brand.

Questions people ask

Why does ChatGPT give me a different revenue number every time I ask?

Usually because each file or connection uses a different revenue basis. Amazon ordered sales, Amazon settlement payouts, Shopify gross sales and Shopify net sales are all called “sales” somewhere. The model uses whichever one it happens to find. Ask it to name the basis of every figure.

Can I fix this with a better prompt?

For one conversation, partly. You can tell the model which periods are closed and which listings are one product. But the next chat, teammate or tool starts from zero. The lasting fix is putting those definitions in the data itself.

Does connecting Amazon or Shopify to Claude through MCP fix the numbers?

No. A raw MCP connector (MCP is the standard way AI apps connect to other software) gives the model fresher rows, not the meaning of the rows. It still has to guess the revenue basis and whether a period is finished. It also has to guess which rows belong to the same product.

Why does Meta show almost no return on ads that send people to Amazon?

The Meta Pixel is tracking code that runs on your own website. An Amazon checkout happens on Amazon, where your pixel is not installed. Meta cannot see those sales. Judge Amazon-bound campaigns on Amazon’s own numbers instead.

How long should I wait before judging last week's Amazon ad results?

At least until the attribution window has closed. That is the time after a click in which Amazon still credits a sale to the ad. It is 7 days after the last click for Sponsored Products in Seller Central, and 14 days for Sponsored Brands and Sponsored Display. Before then, ad sales for that week are still rising.

Is Claude better than ChatGPT at reading ecommerce data?

Both fail in the same ways. The problem is missing definitions in the data, not the model’s arithmetic. A stronger model makes the wrong assumption sound more convincing.