Data race on sequence index
performanceActiveStableA data race exists for the sequence index, which can lead to inconsistent results and crashes during concurrent access.
Score Breakdown
Heuristic ranking from public discussion signals — not a validated prediction of commercial opportunity, demand, or willingness to pay.
Composite 66/100 (High, unvalidated). Top driver: Willingness to pay (30% weight, 22.5 pts).
Heuristic only — often urgency map or random scaffolding on ingest, not measured mention frequency. Maps to XPS relevance (with market size).
LLM/mock judgment of intensity from title/summary text — not ops or ticket data. Maps to XPS quality (with willingness to pay).
LLM/mock purchase-intent guess from text — not invoices, surveys, or paid seats. Maps to XPS quality.
Heuristic/scaffold (often random or fixed on insert) — not a verified mention trajectory. Maps to XPS novelty.
Heuristic/scaffold (often random or fixed) — not TAM research. Maps to XPS relevance (with frequency).
Catalog notes (not predictive analysis)
Data race on sequence index (performance). Catalog heuristic opportunity score: 66/100 — a chosen formula over discussion-signal facets, not evidence of demand, conversion, or willingness to pay. Treat as browsing rank, not a commercial prediction.
A data race exists for the sequence index, which can lead to inconsistent results and crashes during concurrent access.
Source Examples
“Sharing `PySeqIter` across threads use-after-frees the sequence under free-threading # Crash report ### What happened? On a free-threaded build, advancing a single, shared iterator over a legacy sequence (a type implementing `__getitem__`/`sq_item` but not `__iter__`/`tp_iter`) from multiple threads use-after-frees the underlying sequence object. The defect is in the fallback iterator itself, so it affects pure-Python classes and C extension types alike. `list` (whose `list_iterator` was hardened for free-threading in https://github.com/python/cpython/pull/115605) is clean. `iter_iternext()` in `Objects/iterobject.c` has no synchronization: https://github.com/python/cpython/blob/9cbd578e38a5ef59f9462b61d1e03338b208cba6/Objects/iterobject.c#L52-L83 The iterator owns exactly one reference to `it_seq` and releases it exactly once, on exhaustion. With the GIL that invariant holds because only one thread is ever inside `tp_iternext`. Without the GIL: 1. **Double DECREF on exhaustion.** Several threads read the same non-`NULL` `it->it_seq`, all observe out-of-bounds from `PySequence_GetItem`, and each executes the `Py_DECREF(seq)`. The single owned reference is released N times, so the sequence can be freed while other code is still holding refs to it. 2. **Borrowed pointer outliving the object.** Thread A is inside `PySequence_GetItem(seq, ...)` (an arbitrarily long call) and thread B takes the exhaustion path and drops what may be the last reference. A then operates on freed memory. (There's also a data race for `index` that #115605 fixed for lists with atomic stores, but per #124397, that's acceptable.) `calliter_iternext` has the same defect: `it_callable`/`it_sentinel` are cleared with non-atomic `Py_CLEAR` on the exhaustion path and read without protection, so a shared `callable_iterator` can double-release and use-after-free them the same way. There's prior art for this issue class in #154043, #154108, #154130 and the same fixes can be applied here. ### Reproducer ```python import sys, threading THREADS = 16 ROUNDS = 2000 class PySeq: # sq_item only -> iter() falls back to PySeqIter def __init__(self, n): self.n = n def __len__(self): return self.n def __getitem__(self, i): if i >= self.n: raise IndexError(i) return i def drain(it, barrier): barrier.wait() while True: try: next(it) except StopIteration: return except Exception: return # torn it_index can walk off the end; benign def hammer(make): for _ in range(ROUNDS): it = make() barrier = threading.Barrier(THREADS) ws = [ threading.Thread(target=drain, args=(it, barrier)) for _ in range(THREADS) ] for w in ws: w.start() for w in ws: w.join() del it seq = PySeq(4) print("gil enabled:", sys._is_gil_enabled()) print("iter type:", type(iter(seq)).__name__) before = sys.getrefcount(seq) hammer(lambda: iter(seq)) after = sys.getrefcount(seq) print( f"before={before} after={after} -> {'CORRUPTED' if after != before else 'ok'}" ) ``` #### Observed results All on macOS / arm64 (Apple Silicon): | Build | `PySeq` (`PySeqIter`) | `[0, 1, 2, 3]` (`list_iterator`) | | --- | --- | --- | | main (`3.16.0a0`, free-threading debug build) | SIGSEGV | clean, refcount ok | | 3.14.6 free-threaded | SIGSEGV (3/3 runs) | clean (3/3 runs) | | 3.13.8 free-threaded | SIGSEGV (3/3 runs) | not run | | 3.14.6 GIL-enabled | clean, refcount ok | clean | | 3.15.x free-threaded | refcount bad (on some runs; see below) | not run | Depending on allocator timing the failure may show up as refcount corruption, type confusion in unrelated code once the freed memory is reused, or as SIGSEGV, or nothing at all: ``` ~/Desktop $ uv run --python=3.15t --isolated repro.py gil enabled: False iter type: iterator ~/Desktop $ uv run --python=3.15t --isolate”
Competitive Landscape
- Existing solutions are either too expensive or too limited
- Most competitors target enterprise, leaving mid-market underserved
- Community scripts and manual processes are the primary alternative
Recommended Next Steps
- ✓Validate pain intensity with 5-10 target customer interviews
- ✓Build minimal viable solution addressing the core workflow
- ✓Test pricing with early adopters from community forums
Related Pain Points
Target Customers
- IT teams at mid-size organizations (100-2000 employees)
- MSPs and consultants managing multiple client environments
- Teams without dedicated specialist staff for this domain
Monetization Ideas
- 1SaaS subscription model ($99-$499/month depending on scale)
- 2Usage-based pricing aligned with value delivered
- 3Freemium tier to drive adoption and prove value