Every organisation that collects data — which is now nearly every organisation — faces the same structural problem: the data exists, but nobody knows what it means yet. Data science is the field that stands between the raw record and the actionable understanding. The structural pull at the field level is Revelation. Whatever a particular data role looks like day-to-day — a business analyst building dashboards for a marketing team, a data scientist modeling churn for a subscription company, an ML engineer deploying a recommendation system, an analytics manager translating findings for executives who do not speak statistics — the field exists because the data contains things that are hidden until someone makes them visible. The move from hidden to revealed requires someone who can look at a table of numbers and see what is happening underneath them.
Discovery sits tightly alongside Revelation and in some roles — particularly data science and ML research — is arguably the stronger pull. The distinction matters: Revelation is making visible what was already there but buried — the trend in the quarterly numbers that becomes obvious once someone plots it correctly, the customer segment nobody noticed because the reporting structure masked it. Discovery is finding something genuinely new in the data — a pattern nobody expected, a relationship nobody hypothesised. Most working data professionals spend more of their time on Revelation work than on Discovery work. The dashboards, the reports, the "here is what the data says" presentations — these are Revelation. The models, the experiments, the "we found something we weren't looking for" moments — these are Discovery. Both are real; the ratio between them is one of the things that differentiates data roles from each other.
🦊
There's a guide here if you want one
Kitsune can talk through anything on this page — whether it might suit you, what to do next, questions this page doesn't answer. Everything here is yours to read either way.
Statistical-programming barrier
Current fact
The barrier
Data science required dual competency in programming (Python/R/SQL) and statistics — the gateway that kept domain experts out
What changed
Natural-language-to-SQL, AI-assisted scripting, and AutoML lower both barriers simultaneously
Behaviours involved
Domain experts can access data analysis without programming, making domain knowledge the differentiator
Analyzing
Business strategists who ask good questions can bypass analyst intermediary
CommunicatingDiscovering
What this is based on
Power BI Copilot
Tableau natural language query
Amazon Q in QuickSight
Google Gemini in BigQuery
How this could age
Low risk on direction; moderate risk on magnitude
What it does not cover
Collapse is real for descriptive and basic predictive work; causal inference and experimental design barriers intact. Risk: bypassing the barrier also bypasses the statistical understanding that prevents wrong conclusions.
Assessed May 2026
ML deployment barrier
Current fact
The barrier
Building models in notebooks was one skill; deploying to production (containerisation, serving, monitoring) was a different infrastructure-heavy skill
What changed
MLOps platforms, low-code deployment, and AI-assisted infrastructure code compress the DS-to-deployment workflow
Behaviours involved
Data scientists can build and deploy models without dedicated ML engineering support
Building
Researchers can iterate from hypothesis to deployed model without handoff
EngineeringExperimenting
What this is based on
Streamlit
Hugging Face Spaces
SageMaker endpoints
Vertex AI serving
How this could age
Low risk
What it does not cover
Collapse real for standard models; complex high-stakes production systems still require dedicated ML engineering
Assessed May 2026
Domain-crossing barrier
Current fact
The barrier
Data scientists needed years of domain immersion to be effective in a new sector; cross-domain transitions were slow and risky
What changed
LLMs provide rapid domain synthesis; AI-assisted literature review enables fast contextual acquisition
Behaviours involved
Domain transitions feasible in months rather than years
Discovering
Cross-domain pattern transfer faster with AI-assisted context
Pattern-finding
What this is based on
LLM-assisted regulatory mapping
AI-powered clinical trial onboarding
Elicit for domain research synthesis
How this could age
Low risk on direction; moderate risk on overstating barrier collapse
What it does not cover
Deep domain expertise (tacit knowledge, relationships, political sensitivity) still requires time. Surface competence accelerated; depth still earned.
Assessed May 2026
Jobs that did not exist five years ago
These are real jobs that exist now and did not exist before the current wave of AI.
AI Engineer
LinkedIn #1 fastest-growing US job 2024–2025; Onward Search · current fact
low_risk
ML Platform Engineer / MLOps Engineer
Job postings; MLOps ecosystem growth · current fact
low_risk
AI/ML Product Manager
Job postings; PM specialisation trend · current fact
low_risk
Analytics Engineer
dbt ecosystem; data team structure evolution · current fact
low_risk
Prompt Engineer / AI Workflow Designer (DS context)
Vourakis 2026: 1 in 3 AI-related DS postings require RAG/prompt/vector-DB experience · current fact
moderate_risk — title may not survive; the skills will
Bars above the line are the parts of this work that still need a person. Bars below it are what AI can already do. Tap any column to see the actual work behind it.
high ground · holds stronglydeep water · reaches furthest
yours, by strengthAI reach, by depth* depends on the role
For this field we have described what still needs a person, but not scored how strongly. These are the parts named.
Problem framing and question formulationRequires domain knowledge + statistical intuition + judgment about data availability and business relevance. No AI tool currently performs this translation reliably.
Experimental designCausal inference design requires judgment about confounds, validity threats, and appropriate methods that AI assists but cannot drive
Domain-contextual interpretationAccumulated domain expertise (seasonal patterns, data quirks, institutional knowledge) is irreducible
Stakeholder communication under ambiguityPolitical and communicative work in organisational contexts; requires reading the room, managing expectations, and framing tradeoffs for non-technical audiences
Ethical evaluation of model deploymentBias assessment, disparate impact evaluation, and deployment guardrail design involve human judgment with irreducible accountability stakes
Data science is unique among assessed fields: AI capabilities are not external tools being adopted — they are the field's own output. The capability relevance matrix for DS is effectively a self-referential map: the field builds the capabilities that are reshaping it.
How AI is changing the way in
4 ways into this field, and AI is not doing the same thing to each of them.
Data Analyst (entry-level BI / reporting / SQL)Much harder to enter
Natural-language-to-SQL, automated dashboarding, and AI-assisted BI compress the core tasks. Indeed data & analytics postings -39.8%. 365 Data Science: <6% of AI-requiring DS roles are entry-level. The entry analyst who writes SQL and builds dashboards is being displaced by AI-powered BI tools that non-analysts can use directly. Evidence grade: current_fact (posting data) + inference (displacement mechanism).
Data Scientist (junior, model-building focus)Harder to enter
AutoML compresses standard model-building tasks. Junior DS who build standard models (classification, regression, time-series) face competition from AutoML platforms. But the problem-framing and experimental-design components of the DS role remain human, creating a floor for the role that analyst roles lack. Evidence grade: inference.
ML Engineer / AI Engineer (entry-level)Slightly harder to enter
Role is growing — BLS 34% growth projection. Entry-level ML engineer positions exist and are in demand, though the bar has risen (now expected to have MLOps, agentic AI, and deployment skills). The shortage pattern from Cybersecurity (AI extends capacity rather than compresses headcount) partially applies here. Evidence grade: current_fact (BLS projection) + inference.
Analytics Engineer / Data Engineer (entry-level)Slightly harder to enter
Data engineering demand remains strong (infrastructure must be built and maintained). AI tools assist but pipeline architecture, data governance, and system design remain human-intensive; some sub-tasks trend toward moderate as AI-assisted pipeline tooling matures. Evidence grade: inference.
That is everything we currently know about AI in Data Science / Analytics. It shows where things are moving so you can choose which way in suits you.
People drawn to Data Science / Analytics are often drawn to these. Most sit in a different part of the terrain.