PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 17 hours 35 minutes 40 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
2 / 1 024 max
Total
240 522
Rejected
0
llm:64fe130c9b0a4f149b011c667c1079bad96bfb51a595c717244a85dd61e5fab9
Edit
**The categorical columns are almost uniform: no company or ticker dominates, and the most common values each hold only about 0.8% of rows.**
**company_name and stock_name**
- Both columns have 152 distinct values out of 422,729 rows (step-0). The card's "~149" was an approximation.
- They map one-to-one, because each ticker has the same row count as its company (step-0 and step-1). Treat them as one dimension.
- The most common values are ASM International (ASM.AS), BE Semiconductor (BESI.AS) and Soitec (SOI.PA), each with 3,475 rows (0.82%).
- Next comes Melexis NV (MELE.BR) with 3,473 rows.
- Then Elmos Semiconductor (ELG.DE), Infineon (IFX.DE) and SÜSS MicroTec (SMHN.DE), each with 3,448 rows (0.82%).
- Aixtron (AIXA.F) and centrotherm international (CTNK.F) have 3,445 rows (0.81%).
- Weebit Nano (WBT.AX) and BrainChip (BRN.AX) have 3,439 rows (0.81%).
- IQE plc (IQE.L) has 3,431 rows.
- A large group of US names, including NVIDIA, Applied Materials, Analog Devices, Marvell, Wolfspeed, Himax, Axcelis and inTEST, each have 3,415 rows (0.81%) (step-3).
The top of the ranking is essentially a tie. The spread between the most and least frequent of the top 20 is about 60 rows. This reflects how long each stock has been trading and differences in exchange holiday calendars, not real concentration. The European listings (Euronext, Xetra and Frankfurt) have the most rows, and the US and other listings have slightly fewer. Most series start on 2012-01-02 or 2012-01-03 and end on 2025-08-01. The Australian listings (Weebit Nano, BrainChip) run to 2025-08-04 (step-1).
**Pipeline columns**
- `_batch_id` (`batch_d9e1fae7c9b5`), `_source_file` (the Kaggle `semi_conductor_se.csv`) and `_source_system` (`api`) each have a single value covering 100% of rows (step-0). They carry no analytical information.
**Takeaway:** Every company contributes about the same number of daily observations. Row counts can therefore be used as-is in cross-company comparisons without weighting for sample size. The dataset has no other categorical columns, such as sector or exchange. Exchange can only be inferred from the ticker suffix (.AS, .DE, .PA and so on).