Methodology
How this was analyzed
This is an AI-driven corpus analysis, not my retrospective self-assessment. Using an external model reduces the risk of my own confirmation bias, but it does not make the analysis bias-free.
External AI perspective
The model proposed the analytical dimensions, implemented the classifier, inspected ranked evidence, and generated the presentation. The same product-agnostic relation rules are applied regardless of database or employer.
Unit of analysis
One publication can discuss several databases. Each explicit publication × database pair is classified separately and attached to the employer period containing its publication date.
Conservative evidence
Titles and AI-generated summaries are scored automatically, supplemented by body prose that explicitly names a database and contains evaluative or comparison language. Code, diagnostic output, link destinations, generic problem language, negated statements, and reported claims are excluded from sentiment.
Evaluation
Explicit positive, critical, and relational phrases produce a five-point score: strongly critical (−2), critical (−1), neutral or descriptive (0), supportive (+1), and strongly supportive (+2). Comparisons are assigned by semantic role: a product used as the capability benchmark in “A now matches B” supports both products, while plain API compatibility remains neutral.
Criticism is not bashing
Intent and scope remain separate from evaluation. A reproduced optimizer bug is critical but specific and evidence-backed. “This database is bad” would be product-wide. The article reports those cases independently.
Attention is separate
The charts condition on articles that discuss a database. Writing more often about an employer's product changes attention, not necessarily tone, and is not counted as positive sentiment.
Limits
This is descriptive, not causal. Employer and historical period are inseparable; AI-generated summaries and vocabulary rules can carry model bias, miss irony, or lose nuance; and small samples are shown rather than generalized.
The classifications, excerpts, signals, confidence, and aggregates are published in data.json. Regenerate them with python util/build_database_tone.py. The ranked audit is used to inspect both extremes before publishing.