# Analysis and Robustness Plan

**Version:** 1.0  
**Date:** 23 August 2026

## 1. Analytical objective

The study evaluates whether Pakistan experienced six distinct forms of change between 2014 and 2025: cultural production, audience attention, economic sustainability, institutional maturity, social inclusion, and internationalization. These dimensions must be reported separately. The study will not manufacture a single "rise of the industry" score whose weights quietly decide the conclusion before analysis begins.

## 2. Evidentiary status

The current package contains a research design and a starter evidence registry. It does not yet contain the 360-track dataset, artist interviews, contracts, income observations, or complete historical platform series. Every proposition is therefore provisional.

## 3. Samples

- **Track sample:** 360 tracks, 30 from each complete release year, 2014-2025.
- **Artist sample:** 80-100 artists and acts linked to the track sample plus purposive negative and comparison cases.
- **Content sample:** 120 tracks/videos, 10 per year.
- **Interviews:** target 38 across artists, production, intermediation, brands/platforms, live music, rights/policy, journalism, and audiences.
- **Epilogue:** January-August 2026, descriptive only.

The track sample is diversity-stratified and purposive. It supports structured comparison and mechanism analysis, not a claim that 30 tracks statistically represent every release in a year.

## 4. Measurement architecture

### Cultural production

Indicators: candidate-universe release count, sampled genre and language diversity, collaboration rate, independent entry, production route, and producer network breadth.

### Audience attention

Indicators: platform-specific view/stream observations, chart entry, peak rank, chart duration, search interest, short-form reuse, and attention persistence. Normalize only within a platform and comparable time window.

### Economic sustainability

Indicators: share of personal income from music, number and diversity of revenue sources, payment delays, unpaid work, documented fees or royalties, reliance on outside work/family support, and expected three-year career continuity.

### Institutional maturity

Indicators: written contracts, master and publishing ownership, metadata quality, label/distributor/publisher/manager relationships, rights administration, venue and festival continuity, financing sources, and dispute-resolution pathways.

### Social inclusion

Indicators: representation by gender, language, region, city, and class-access proxy, followed by conversion outcomes such as ownership, payment, lead credit, decision authority, and career continuity.

### Internationalization

Indicators: foreign chart entry, audience geography, diaspora collaboration, touring, international label/distributor relationships, awards, sync placements, and revenue geography where available.

## 5. Descriptive analysis

1. Produce annual and three-year rolling summaries for every dimension.
2. Report medians and interquartile ranges for skewed attention variables.
3. Keep platform measures separate; create indexed within-platform series only after checking metric continuity.
4. Show missingness and deleted-content rates beside every trend.
5. Compare release cohorts at equal exposure windows where possible.
6. Report the number of independent observations behind every percentage.

## 6. Concentration analysis

Within the observed dataset, calculate:

- Top 1, top 5, and top 10 shares of platform attention.
- Herfindahl-Hirschman concentration by artist, channel, brand, label/distributor, producer, and city.
- Gini coefficients for attention and documented income where data density permits.
- Share of artists receiving 50 percent of observed attention.

Because the sample deliberately includes high-attention tracks, these estimates describe concentration within the research sample unless a broader candidate-universe dataset is constructed. Do not relabel them national market shares.

## 7. Network analysis

Build two-mode and projected networks connecting artists to producers, labels, distributors, brands, programmes, and collaborators. Examine:

- Degree and betweenness centrality.
- Recurring institutional brokers.
- Regional and gender homophily.
- Entry pathways into central networks.
- Whether one-off branded exposure produces subsequent independent ties.

Run sensitivity tests with and without large rotating programmes such as Coke Studio so one institutional hub does not turn the entire graph into a logo.

## 8. Career continuity and labour analysis

For artists with reliable career histories:

- Define entry as first documented public release.
- Define activity by release, live, production, or other paid music work within a 12-month window.
- Use Kaplan-Meier-style descriptive survival curves for time to a 24-month inactivity gap.
- Treat re-entry as a separate event rather than permanent failure.
- Compare pathways by institutional support, gender configuration, origin, and income dependence.

The analysis is descriptive because entry dates, inactive periods, and unobserved work may be measured imperfectly.

## 9. Exploratory models

Use models only after descriptive and missingness review.

### Attention model

Outcome: log of platform-specific attention rate or chart success. Predictors: release year, exposure age, stratum, language, genre, institutional support, collaboration structure, video presence, artist prior audience, and promotional evidence.

### International breakthrough model

Outcome: foreign chart entry, documented overseas audience threshold, or international award/placement. Predictors: diaspora ties, shared-language market, distributor/label relationship, brand platform, collaboration, and genre.

### Career-sustainability model

Outcome: majority income from music, three-year continuity expectation, or observed active-career duration. Predictors: revenue diversification, contract status, ownership, live work, platform reach, class-access proxy, gender, geography, and institutional support.

Use robust uncertainty estimates and report predictive associations, not causal effects. With a purposive sample, p-values do not turn convenience into representativeness.

## 10. Qualitative analysis

- Code interviews using a combined deductive and inductive framework.
- Start with access, costs, income, rights, platform power, sponsorship, live infrastructure, class, gender, region, language, diaspora, regulation, and definitions of industry.
- Add emergent codes through a logged codebook revision process.
- Write within-case memos before cross-case comparison.
- Search deliberately for negative cases that contradict each proposition.
- Compare public narratives with private mechanisms without exposing confidential sources.

## 11. Content analysis

For the 120-item corpus:

- Track theme and visual-code prevalence by four periods: 2014-2016, 2017-2019, 2020-2022, and 2023-2025.
- Compare branded, label-supported, and independent pathways.
- Compare regional-language visibility with credit, ownership, and follow-on outcomes.
- Examine whether representations of aspiration, class, gender, nation, religion, and region change with platform and genre.
- Report intercoder agreement and translation uncertainty.

## 12. Process tracing

Investigate mechanisms around documented turning points:

- April 2014 mobile spectrum auction.
- 2015 local streaming launches.
- January 2016 YouTube restoration.
- 2019 local-platform payment disputes.
- 2020 pandemic disruption and home-production acceleration.
- February 2021 Spotify entry.
- 2022 international breakthrough cases.
- 2023 national music-policy process.
- 2024 global-label and branded-distribution partnerships.
- 2025 digital-regulation changes.

For each event, specify the expected mechanism, observable implications, beneficiaries, excluded groups, and alternative explanations. Temporal sequence alone is not causal identification.

## 13. Robustness and sensitivity checks

| Risk | Required check |
|---|---|
| Release-age bias | Equal exposure windows; views per month; cohort-matched comparison |
| Platform metric drift | Record definitions and interface changes; avoid splicing incompatible series |
| Deleted/private content | Missingness category, archive evidence, and sensitivity excluding affected years |
| Paid promotion or bots | Flag campaign evidence, suspicious discontinuities, and platform-adjusted chart data |
| Survivor bias | Include inactive artists, failed ventures, rejected applicants, and low-attention tracks |
| Celebrity bias | Cap high-profile interviews; recruit by sampling cells rather than referrals alone |
| Cross-platform incomparability | Analyze separately; only index after within-platform standardization |
| Currency and inflation | Preserve nominal PKR; report conversion date; use CPI-adjusted PKR for time comparison |
| Geography ambiguity | Separate artist origin, residence, production location, and audience geography |
| Language ambiguity | Multi-label coding, translator notes, and sensitivity excluding uncertain items |
| Missing income | Report missingness; use bands; do not impute exact earnings from public streams |
| 2026 incompleteness | Exclude from annual trend and pooled complete-year models |
| Brand self-reporting | Triangulate first-party claims with channel data, contracts, and participant accounts |
| Researcher network bias | Publish recruitment routes and compare reached versus missing sampling cells |

## 14. Triangulation rules

Classify a claim as:

- **Established:** supported by a primary record or at least two independent high-quality sources with no material contradiction.
- **Strongly supported:** consistent across multiple source types, including direct participant or documentary evidence.
- **Plausible:** supported by limited cases or theory but lacking denominators or independent confirmation.
- **Contested:** credible sources materially disagree.
- **Unknown:** evidence is insufficient.

A first-party platform or brand claim can establish what the organization announced. It cannot, by itself, establish the market-wide effect of that announcement.

## 15. Decision rules for propositions

For each proposition, publish:

1. Supporting evidence.
2. Contradictory and negative cases.
3. Coverage across years, genres, regions, and participant types.
4. Evidence-quality rating.
5. Final status: supported, partly supported, unsupported, or unresolved.
6. The observation that would most likely change the conclusion.

## 16. Reproducibility outputs

Release, subject to rights and ethics:

- Source map and capture log.
- Sampling and exclusion log.
- Dataset schemas and de-identified analytic tables.
- Code used for cleaning, charts, and models.
- Interview and content codebooks.
- Reliability and missingness reports.
- A limitations register documenting inaccessible data.

Do not release participant keys, confidential transcripts, copyrighted media, platform credentials, or private contracts.
