PHPMem v2.0.1
Version
1.6.45
Uptime
18 days 8 minutes 20 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
4 / 1 024 max
Total
242 724
Rejected
0
llm:2056b7a8665f63e04fec24815284fc3d08d8cb5a409dba4edbb1db2cafb0ae34
Edit
# Three analytical questions this dataset could answer
The investigation was cancelled before I ran any comparison queries. Every query I completed returned the same dataset-wide summary, so none of the three questions below is answered yet. The evidence shows what the data can support, not what it finds.
## What the data contains (supported by the queries)
- **Coverage:** 422,729 daily rows for 152 companies and 152 tickers, from 2012-01-02 to 2025-08-04.
- **Post-ChatGPT share:** 99,984 rows (23.7%) fall after the ChatGPT launch. 96,711 rows are from 2023 onward.
- **Price level:** median close is about 35.57.
- **Intraday volatility:** average daily range is 3.69% of close.
- **Volume is heavily skewed:** median volume is 898,500, mean is about 8.47M, and the maximum is about 3.69B.
- **Zero-volume rows:** 4,888 rows, about 1.2% of the total.
- **Completeness:** close and volume have no nulls (422,729 non-null each).
## The three questions
### 1. Did volatility and trading activity change after ChatGPT's launch, and did it differ by company?
- **Why it fits:** The data spans about 10.9 years before the launch and about 2.7 years after, with 99,984 post-launch rows. That is enough for a before/after comparison of average daily range (3.69% overall) and volume.
- **Not yet known:** I have no pre/post split, so I can't say whether anything changed.
- **Caveat:** A sharper version would compare AI-exposed names with the rest. That needs a sector or AI-exposure mapping, and I haven't confirmed one exists in the data. Other events, such as rate hikes, would also confound a simple before/after comparison.
### 2. What drives the extreme skew in trading volume, and how much of it is concentrated in a few stocks or days?
- **Why it fits:** The mean (8.47M) is about 9.4 times the median (898,500), and the maximum is 3.69B. A small number of tickers or event days likely dominate the average.
- **Not yet known:** Which tickers or dates produce the spikes. This would also show whether the maximum is a real event or a data or split artifact.
### 3. Are the 4,888 zero-volume rows a data-quality problem or a real liquidity signal?
- **Why it fits:** Zero volume on about 1.2% of rows is unusual for actively traded stocks.
- **Not yet known:** Whether these rows cluster in particular tickers, early years, or halts and delistings.
- **Why it matters:** The answer affects every volume-based metric, and it could reveal that the panel is unbalanced. The average is about 2,781 rows per company. That is fewer than a full 2012–2025 history would give, which suggests some companies have shorter histories, though I haven't checked this directly.
## Bottom line
The dataset is clean on nulls and large enough for event-study, distributional, and data-quality questions. Answering any of them needs grouped queries by period, ticker, or date, and none were completed.