Case Study · The Growth Journey
Customer & Competitive Intelligence System — Amazon Reviews V1
From 2,302 Amazon reviews to competitive evidence across five premium audio brands.
Originally developed as the final project of a Data Analytics & AI Bootcamp and later refined as a portfolio project and proof of work.
Case Study — Data & Market Intelligence
Customer & Competitive IntelligenceAmazon Reviews — V1 2,302 Amazon reviews. Five premium audio brands. Seven product dimensions.
Project at a glance
The problem started before the technology
The idea emerged while I was looking for a question for my final Data Analytics & AI project. During that process, Amazon’s reports on the growth of independent businesses selling through its marketplace caught my attention.
In 2023, Amazon reported that more than 10,000 independent sellers had surpassed $1 million in sales for the first time. By 2024, more than 55,000 sellers had generated over $1 million, and in 2025 the figure exceeded 75,000, up 36% from the previous year. The 2023 figure specifically measures sellers crossing that threshold for the first time, so it is not directly equivalent to the total figures reported for 2024 and 2025.
More than the figure itself, I was interested in the business reality behind it: thousands of companies competing in an environment where products, prices, competitors and customer opinions can be observed almost side by side.
On Amazon, a buyer can compare features and alternatives and, a few centimetres further down the page, find hundreds or thousands of people explaining what they value, what disappointed them, what worked and what they would not buy again. A single review may be anecdotal. Thousands of them contain patterns, frictions and comparisons that can potentially say something about a product’s position relative to its competitors.
The project initially began as an exercise in sentiment analysis, but it soon evolved into a broader question.
If a brand competes in this segment, where is it gaining and where is it losing against its direct competitors according to customer reviews, and which of those differences can be supported by enough evidence to amount to more than a superficial comparison of averages?
To work with a concrete market, I chose premium over-ear active-noise-cancelling headphones. Sennheiser became the anchor brand against Bose, Sony, Bang & Olufsen and Bowers & Wilkins. Using a real company and market made the question more concrete; it did not, however, turn an academic portfolio project into client work that never took place.
Build the system before looking for answers
Answering the question required more than downloading reviews and calculating averages. I needed to build a process that preserved the connection between what customers wrote and the conclusions that would appear at the end of the analysis.
That meant solving three different problems: defining which products actually belonged in the competitive set, transforming unstructured language into analysable variables, and verifying that the model used to interpret the reviews was sufficiently reliable before applying it to the full dataset.
I organised the project around the six phases of CRISP-DM and a bronze/silver/gold data architecture. Python was used for exploration, preparation and analysis; MySQL to structure and validate information; language models to interpret reviews; statistical methods to evaluate models and test selected findings; and Tableau to build the visual analysis experience.
The source was the public Amazon Reviews 2023dataset, developed by McAuley Lab at UC San Diego. I worked within the Electronics category, selected specific product lines in the premium ANC segment and set a minimum price of $150. The initial dataset contained 2,312 reviews; after cleaning and removing ten strict duplicates, 2,302 observations remained for analysis.
Reaching that dataset produced one of the project’s first important lessons. An initial filter classified some AKG products as Sennheiser, while an overly generic search returned 1,154 matches for Sony. The problem was not yet AI or statistics. It was that I had not defined precisely enough what I intended to compare.
The solution was to work with specific product families — Momentum, QuietComfort, WH-1000X and the corresponding lines from the other brands. The unit of analysis also became more precise: I was not comparing “Sennheiser versus Bose” in absolute terms, but selected premium ANC product lines within the scope supported by the dataset.
That correction changed less code than meaning. An analysis can be technically sound and still answer the wrong question if the reality entering it is poorly defined.
The system ultimately evaluated seven dimensions: sound quality, noise cancelling, comfort, battery, build and design, value for money, and software/connectivity. The last of these was not part of the initial schema; it emerged during manual labelling, when references to Bluetooth, firmware and apps showed that the original conceptual model was leaving out a relevant part of the product experience.
Build evidence before scaling AI
Connecting a model and processing all 2,302 reviews directly was technically possible. The problem was that a well-structured response does not by itself prove that the interpretation is correct. If a systematic error existed, scaling the workflow would only reproduce it thousands of times.
That is why I first built a ground-truth set of 200 manually labelled reviews, 40 per brand. Each review received an overall sentiment label and labels for the seven product dimensions using the same schema that the models would later apply.
I then compared Llama 3.3 70B and GPT-OSS 120B. Before seeing their results, I defined the criteria they had to meet. Both models passed the standard, and the paired Wilcoxon comparison produced p = 0.757, so the data did not support declaring a statistical winner.
Two models met the standard. Choosing between them required more than looking for a winner.
| Metric | Llama 3.3 70B | GPT-OSS 120B |
|---|---|---|
| Polarity accuracy | 90.0% | 91.5% |
| Cohen’s Kappa | 0.722 | 0.763 |
| Macro F1 | 0.763 | 0.608 |
| Schema retries | 3.5% | 5.5% |
| Time / review | 9.4 s | 28.8 s |
| Estimated total cost | $0.41 | $0.34 |
Selected for its combination of speed, operational reliability and more balanced behaviour on selected metrics, while preserving the advantages observed for GPT-OSS on other measures.
What the analysis found
Once the reviews had been processed, the challenge was no longer mainly technical. The question became what the observed differences actually meant.
Averages made it possible to compare the five brands quickly across the seven dimensions. But an observed difference was not automatically a demonstrated competitive advantage. For selected headline findings, I used bootstrap to estimate confidence intervals and test whether the differences held up under additional scrutiny.
For sound quality, Sennheiser had the highest observed average at 0.755 versus Bose at 0.748. However, the confidence interval for the difference crossed zero. For noise cancelling, Sennheiser scored 0.299 versus Bose at 0.781 and Sony at 0.767, and the gaps against both brands remained statistically significant. For battery, the gap versus Sony was also statistically supported.
Observing a difference does not mean demonstrating an advantage.
Sound quality
Sennheiser 0.755
Bose 0.748
95% CI of the difference
−0.071 → +0.088
Highest observed average. No statistically demonstrated edge.
Noise cancelling
Sennheiser 0.299
Bose 0.781
Sony 0.767
Statistically significant gap vs Bose
Statistically significant gap vs Sony
Supported disadvantage versus Bose and Sony within this sample.
Battery
Sennheiser 0.408
Sony 0.741
Statistically significant gap vs Sony
Supported disadvantage versus Sony within this sample.
“Has the highest average” does not necessarily mean “has a demonstrated advantage.” And “ranks last” does not automatically mean “is significantly worse than every competitor.”
Competitive position changes by dimension.
The same brand can occupy very different positions depending on which part of the product experience is being examined.
The analysis does not produce a single brand hierarchy that holds across the entire product experience.
When evidence changed the answer
The analysis did more than produce information about the brands. It also forced me to revisit some of the hypotheses I had started with. Sound quality was especially important because it required abandoning a favourable interpretation: “Sennheiser wins on sound” was a narrative consistent with the brand’s reputation and easy to communicate; after statistical testing, it was no longer a defensible conclusion.
No statistically demonstrated edge.
The standard had to remain the same whether the evidence supported the story or contradicted it.
The project also included a small experimental layer designed to detect potential safety signals that might disappear inside averages.
Brands were anonymised in the public dashboard because a few individual cases do not provide a responsible basis for broad claims. The detector also could not estimate how many cases might have used different language and remained unidentified. The three findings were therefore observed signals, not a failure rate or an estimate of prevalence.
From finding to decision
Finding a competitive difference does not automatically resolve what an organisation should do about it. The noise-cancelling gap versus Bose and Sony could justify further investigation. It did not, however, provide enough evidence on its own to recommend a costly or difficult-to-reverse decision such as redesigning hardware or changing a product platform.
Why does this gap appear versus Bose and Sony, and do we see the same pattern in our internal data?
The next layer of evidence might come from returns, warranty claims, customer support, proprietary surveys, telemetry or additional research using a different design.
Reviews can help indicate where to look. They do not necessarily carry enough authority to decide what to do.
The greater the cost, risk or irreversibility of a decision, the stronger the evidence should be before committing to it. In that sense, analysis does not replace professional judgment; it improves the context in which that judgment must be exercised.
Where the conclusions end
The project is useful only if its limits remain visible. The 2,302 reviews allow patterns to be observed within the analysed dataset, but they do not represent the entire customer experience of these brands or a controlled study of the whole premium audio market.
- Source
- Amazon only; data through September 2023.
- Sampling
- 500-review cap for four brands; Bang & Olufsen remained at 312, the full population available under the filters.
- Product generations
- Some product lines group different generations within the same family scope.
- Price coverage
- Sennheiser: approximately 47.3% price coverage.
- Ground truth
- Single human annotator; inter-annotator agreement was not measured.
- Safety detection
- The keyword-based detector cannot estimate the false-negative rate.
- Reuse
- Architecture designed for reuse; cross-category transfer has not yet been demonstrated.
The architecture was designed for reuse. Whether it maintains its quality when moved to another category still needs to be demonstrated.
Project evidence
Project evidence and outputs
The case study explains the problem, the decisions and the limits. External resources allow the work to be explored in greater depth.
Explore the full analysis
Rankings, product dimensions, model evaluation, mention rates and additional levels of detail are available in the interactive dashboard.
Review the public technical evidence
The public repository includes selected exploration, extraction, SQL and project documentation. The complete analytical engine remains in the private repository.
Next iteration
Amazon Reviews V2: test before generalising
The next iteration will remain within the Amazon Reviews line, but it will change both the product category and the anchor brand. Both are still to be decided.
The purpose of V2 is not to confirm that V1 works universally. It is to test what happens when the architecture faces a different context: different customer language, different product dimensions and a different competitive logic.
It will also incorporate improvements that emerged directly from V1, including replacing the keyword-based safety detector with a contextual evaluation integrated into each review and reorganising the dashboard so that the summary layer is more clearly separated from detailed exploration.
V2 should not prove that V1 was right. It should test which parts of V1 deserve to survive.
There is a later and separate direction: exploring whether the logic can evolve beyond Amazon to sources such as Google Maps, Yelp or TripAdvisor. That stage would introduce new problems of data acquisition, normalisation and comparability that have not yet been designed. It is not part of V2 and does not represent a capability that V1 has already demonstrated.
What remains after 2,302 reviews
The project began with a question about Amazon, reviews and AI and ended by forcing me to pay as much attention to the limits of a conclusion as to the conclusion itself.
Some filters had to be corrected, one analytical dimension emerged only after manual labelling had begun, two models met the standard without producing a clear statistical winner, and a favourable conclusion on sound quality stopped being defensible once it was subjected to greater scrutiny.
Technology made it possible to analyse information at a scale that would have been difficult to handle manually. Statistics helped distinguish some signals from differences that could be explained by variation. But neither replaced the need to define the problem correctly, decide what evidence was sufficient, recognise when a hypothesis had failed, or prevent a finding from becoming a recommendation too early.
The value of the project is not only in finding differences between five brands. It is in building a process capable of distinguishing between what appeared to be true, what the data suggested, and what the evidence actually allowed me to defend.
V1 is useful as a starting point not because it closed every question, but because it made sufficiently clear which questions deserve to be tested again.
Project note. Customer & Competitive Intelligence System — Amazon Reviews V1 was originally developed as the final project of a Data Analytics & AI Bootcamp and later refined as a portfolio project and proof of work. Sennheiser was used as the anchor brand; Sennheiser, Bose, Sony, Bang & Olufsen and Bowers & Wilkins did not commission or participate in the analysis. There is no commercial relationship between this project and Amazon. The analysed data comes from the public Amazon Reviews 2023 dataset from McAuley Lab, UC San Diego.

