← Back to history

Pipeline run

ecbfac28-c2ec-4008-8a24-31eff49881c6

Pipeline LLM cost (USD)
API 1: $0.0031 API 2: $0.0000 API 3: $0.0000 Total: $0.0031

Client output enrichment

v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA description
Nature of work
—
no_db_connection
Tech stack maturity
Mainstream Modern
AI index (0 = no AI use, 5 = totally AI-dependent · v2.1)
1.70 / 5
· Title match
✓ Has AI skill
✓ AI skill (primary)
· AI skill (secondary)
· On AI team
· Builds AI products
vocab breakdown (legacy)
Assistants (×1): —
Frameworks (×2): —
Models / concepts (×3): AI, Generative AI, Machine Learning
Evidence — skills matched in JD (17)
Python SQL Machine Learning Databricks Apache Spark Natural Language Processing Microsoft Power BI Tableau Statistical Modeling Data Visualization ETL Big Data Experimental Design AWS Microsoft Azure Generative AI Git
Skill cluster (0 dimension groups, role-scoped)
No dimension groups computed for this JD.
Status: jd_deconstruction_v4_done Created: 2026-09-05T14:27:54.439652Z Updated: 2026-09-06T05:27:15.576464Z Run duration: 10539 ms
Flow v4 catalog deconstruction skill_library_v4

1 LLM mention extraction + deterministic catalog sweep

2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)

3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim

Role Chosen role & resolution

Data Scientist

in_db body_evidence family · data-scientist

slug: data-scientist · role_id: 293411ee-d03b-59f7-a4e0-60756ab78387 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)

The role is for a data scientist with either an advanced degree or a bachelor's degree plus two years of relevant experience. The core technology stack includes Python, SQL, cloud platforms like AWS or Azure, and tools such as Databricks and Spark. Distinctive requirements include expertise in machine learning algorithms and experience with data visualization tools like Power BI or Tableau. The position involves working in cross-functional teams to deliver data-driven solutions, with a focus on statistical modeling and data engineering.

Job description

REQUIREMENTS:
• Advanced degree with no prior experience required, or Bachelor's degree plus 2 years of experience in applied data science or related role
• Preferred Fields of Study: Data Science, Computer Science, Statistics, Mathematics, Engineering, or related quantitative field
• Additional relevant certifications in machine learning, AI, or data science a plus
• Strong proficiency in Python, SQL, and data analysis tools
• Expertise in machine learning algorithms, statistical modeling, and data visualization
• Experience with cloud platforms (AWS, Azure) and modern data architectures
• Proficient in data engineering and ETL processes
• Experience with Databricks, Spark, and big data technologies
• Knowledge of generative AI and natural language processing
• Strong understanding of statistical concepts and experimental design
• Experience with data visualization tools (Power BI, Tableau)
• Excellent problem-solving and analytical skills
• Strong communication skills and ability to present complex findings to stakeholders
• Experience working in cross-functional matrix environments
• Demonstrated success in delivering data-driven solutions
• Ability to work independently and collaborate effectively with teams
• Experience with version control systems (e.g., Git)
• Travel may be required (0-25%)

Skills from this JD

L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.

L0 · Anchor 4 skill(s)
Python Programming Languages Python and SQL universally; R in research-leaning and life-sciences teams.
SQL Programming Languages Python and SQL universally; R in research-leaning and life-sciences teams.
Statistical Modeling Statistical Modeling & Inference This is the discipline the title names.
Experimental Design Experimentation & Causal Inference The single clearest separator from every neighbouring role.
L1 · Primary 9 skill(s)
Machine Learning Applied ML Modeling & Methods Data Scientists build models constantly — for explanation and forecasting rather than for a production endpoint.
Databricks Data Processing & Pipeline Frameworks pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.
Apache Spark Data Processing & Pipeline Frameworks pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.
Natural Language Processing Applied ML Modeling & Methods Data Scientists build models constantly — for explanation and forecasting rather than for a production endpoint.
Microsoft Power BI Data Visualisation & Reporting Charts and notebooks are the working surface, not a presentation afterthought.
Tableau Data Visualisation & Reporting Charts and notebooks are the working surface, not a presentation afterthought.
Data Visualization Data Visualisation & Reporting Charts and notebooks are the working surface, not a presentation afterthought.
ETL Data Processing & Pipeline Frameworks pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.
Big Data Data Processing & Pipeline Frameworks pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.
L2 · Environment 3 skill(s)
AWS Cloud Platforms Notebooks, warehouses and compute increasingly live in a managed environment.
Microsoft Azure Cloud Platforms Notebooks, warehouses and compute increasingly live in a managed environment.
Generative AI Foundation Model Platforms & APIs LLMs have become a routine analysis tool — for coding assistance, unstructured text coding and summarisation.
L3 · Hygiene 1 skill(s)
Git Version Control Hygiene, and the one engineering habit most worth checking in an otherwise research-leaning candidate.
Consider adding canon exemplars the JD missed

A/B Testing — Experimentation & Causal Inference (L0)

R — Programming Languages (L0)

Hypothesis Testing — Statistical Modeling & Inference (L0)

Data Storytelling — Analytical Communication & Decision Support (L1)

Gradient Boosting — Applied ML Modeling & Methods (L1)

PySpark — Data Processing & Pipeline Frameworks (L1)

Unmapped not resolved, not classified — quarantined for review
AIdata science

Library artifacts (this run)

No artifact rows for this run.
nano JD Parser — gpt-4.1-nano click to toggle
JD type fail
Show raw JSON
{
  "JD_type": "fail"
}
API 1 — extract-from-jd click to toggle
{
  "catalog": "v4",
  "consider_adding": [
    {
      "dimension": "Experimentation \u0026 Causal Inference",
      "reason": "In the role canon\u0027s Experimentation \u0026 Causal Inference - adding it narrows the candidate pool.",
      "skill": "A/B Testing",
      "tier": "L0"
    },
    {
      "dimension": "Programming Languages",
      "reason": "In the role canon\u0027s Programming Languages - adding it narrows the candidate pool.",
      "skill": "R",
      "tier": "L0"
    },
    {
      "dimension": "Statistical Modeling \u0026 Inference",
      "reason": "In the role canon\u0027s Statistical Modeling \u0026 Inference - adding it narrows the candidate pool.",
      "skill": "Hypothesis Testing",
      "tier": "L0"
    },
    {
      "dimension": "Analytical Communication \u0026 Decision Support",
      "reason": "In the role canon\u0027s Analytical Communication \u0026 Decision Support - adding it narrows the candidate pool.",
      "skill": "Data Storytelling",
      "tier": "L1"
    },
    {
      "dimension": "Applied ML Modeling \u0026 Methods",
      "reason": "In the role canon\u0027s Applied ML Modeling \u0026 Methods - adding it narrows the candidate pool.",
      "skill": "Gradient Boosting",
      "tier": "L1"
    },
    {
      "dimension": "Data Processing \u0026 Pipeline Frameworks",
      "reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
      "skill": "PySpark",
      "tier": "L1"
    }
  ],
  "final_skills": [
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams.",
      "skill_id": "646cb350-7c71-5ed2-971e-d141c566ce6e",
      "skill_name": "Python"
    },
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams.",
      "skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
      "skill_name": "SQL"
    },
    {
      "dimension": {
        "display_name": "Applied ML Modeling \u0026 Methods",
        "slug": "applied-ml-modeling-methods"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Data Scientists build models constantly \u2014 for explanation and forecasting rather than for a production endpoint.",
      "skill_id": "28390c4c-2527-5a27-aefe-77206f9a1935",
      "skill_name": "Machine Learning"
    },
    {
      "dimension": {
        "display_name": "Cloud Platforms",
        "slug": "cloud-platforms"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment.",
      "skill_id": "78c5a4b9-2ab5-5279-a688-46f4c3767c3f",
      "skill_name": "AWS"
    },
    {
      "dimension": {
        "display_name": "Cloud Platforms",
        "slug": "cloud-platforms"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment.",
      "skill_id": "060b1d39-d4d8-581d-9d96-65cd45d7cea4",
      "skill_name": "Microsoft Azure"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.",
      "skill_id": "f621b9a9-434a-5aae-a79f-5d858d66973d",
      "skill_name": "Databricks"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.",
      "skill_id": "f885334e-9062-5967-bc62-b5b5a27664cf",
      "skill_name": "Apache Spark"
    },
    {
      "dimension": {
        "display_name": "Foundation Model Platforms \u0026 APIs",
        "slug": "foundation-model-platforms"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "LLMs have become a routine analysis tool \u2014 for coding assistance, unstructured text coding and summarisation.",
      "skill_id": "8e5aece7-c8c5-59a8-b693-4a5df1c2c4ef",
      "skill_name": "Generative AI"
    },
    {
      "dimension": {
        "display_name": "Applied ML Modeling \u0026 Methods",
        "slug": "applied-ml-modeling-methods"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Data Scientists build models constantly \u2014 for explanation and forecasting rather than for a production endpoint.",
      "skill_id": "fdd8da3b-6a74-50f5-83d9-e9b1caabcbb4",
      "skill_name": "Natural Language Processing"
    },
    {
      "dimension": {
        "display_name": "Data Visualisation \u0026 Reporting",
        "slug": "data-visualization-reporting"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Charts and notebooks are the working surface, not a presentation afterthought.",
      "skill_id": "441f4640-fee5-5dfd-aa1a-b3213839ffef",
      "skill_name": "Microsoft Power BI"
    },
    {
      "dimension": {
        "display_name": "Data Visualisation \u0026 Reporting",
        "slug": "data-visualization-reporting"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Charts and notebooks are the working surface, not a presentation afterthought.",
      "skill_id": "b5b389ea-14af-51b7-a7db-416318e9719b",
      "skill_name": "Tableau"
    },
    {
      "dimension": {
        "display_name": "Version Control",
        "slug": "version-control"
      },
      "is_primary": false,
      "layer": "L3",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Hygiene, and the one engineering habit most worth checking in an otherwise research-leaning candidate.",
      "skill_id": "69a9e960-a4b1-5e7e-b55c-04980465e177",
      "skill_name": "Git"
    },
    {
      "dimension": {
        "display_name": "Statistical Modeling \u0026 Inference",
        "slug": "statistical-inference-methods"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "This is the discipline the title names.",
      "skill_id": "8aab31b5-ff5a-5dba-aa99-39b3ae14f3a7",
      "skill_name": "Statistical Modeling"
    },
    {
      "dimension": {
        "display_name": "Data Visualisation \u0026 Reporting",
        "slug": "data-visualization-reporting"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Charts and notebooks are the working surface, not a presentation afterthought.",
      "skill_id": "726b437e-ac64-5fdc-aba9-755c5695974e",
      "skill_name": "Data Visualization"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.",
      "skill_id": "16a8d754-3339-58e3-978c-7e033bb9b94d",
      "skill_name": "ETL"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.",
      "skill_id": "dbe274b4-5b86-599d-bac5-1fc85691053a",
      "skill_name": "Big Data"
    },
    {
      "dimension": {
        "display_name": "Experimentation \u0026 Causal Inference",
        "slug": "experimentation-causal-inference"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "The single clearest separator from every neighbouring role.",
      "skill_id": "2e0ec19d-e2c6-5626-8bc7-667bc7708b08",
      "skill_name": "Experimental Design"
    }
  ],
  "history_run_id": null,
  "jd_parameters": {
    "certifications": [],
    "clientDetails": null,
    "company": null,
    "companySize": null,
    "ctc": {
      "currency": null,
      "max": null,
      "min": null,
      "period": null,
      "raw": null
    },
    "educationRequirements": [],
    "experience": {
      "max": null,
      "min": null,
      "raw": null
    },
    "industryDomain": null,
    "knockouts": {
      "certifications": [],
      "ctc": {
        "currency": null,
        "max": null,
        "min": null
      },
      "educationRequirements": [],
      "experience": {
        "max": null,
        "min": null
      },
      "location": [],
      "noticePeriod": null
    },
    "locations": [],
    "noticePeriod": null,
    "openToRelocate": false,
    "role": null,
    "roleSynonyms": []
  },
  "jd_summary": "The role is for a data scientist with either an advanced degree or a bachelor\u0027s degree plus two years of relevant experience. The core technology stack includes Python, SQL, cloud platforms like AWS or Azure, and tools such as Databricks and Spark. Distinctive requirements include expertise in machine learning algorithms and experience with data visualization tools like Power BI or Tableau. The position involves working in cross-functional teams to deliver data-driven solutions, with a focus on statistical modeling and data engineering.",
  "layer_conflicts": [],
  "nano_parsed": {
    "JD_type": "fail"
  },
  "pipeline": "v4",
  "rejected": false,
  "rejection_code": null,
  "rejection_reason": null,
  "role": {
    "canonical_name": "Data Scientist",
    "family": "data-scientist",
    "match_method": "body_evidence",
    "resolution": "in_db",
    "role_id": "293411ee-d03b-59f7-a4e0-60756ab78387",
    "similarity": null,
    "slug": "data-scientist"
  },
  "role_coverage": [
    {
      "is_jd_role": true,
      "name": "Data Scientist",
      "pct": 0.81,
      "role": "data-scientist",
      "share": 73.4
    },
    {
      "is_jd_role": false,
      "name": "Data Engineer",
      "pct": 0.48,
      "role": "data-engineer",
      "share": 16.9
    },
    {
      "is_jd_role": false,
      "name": "ML Engineer",
      "pct": 0.48,
      "role": "ml-engineer",
      "share": 9.7
    }
  ],
  "run_id": "jdv4-4bea5ac26938",
  "secondary_meta": [],
  "secondary_skills": [],
  "skill_layers": [
    {
      "label": "Anchor",
      "layer": "L0",
      "skills": [
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "Python",
          "origin": "catalog",
          "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams."
        },
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "SQL",
          "origin": "catalog",
          "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams."
        },
        {
          "dimension": {
            "display_name": "Statistical Modeling \u0026 Inference",
            "slug": "statistical-inference-methods"
          },
          "name": "Statistical Modeling",
          "origin": "catalog",
          "rationale": "This is the discipline the title names."
        },
        {
          "dimension": {
            "display_name": "Experimentation \u0026 Causal Inference",
            "slug": "experimentation-causal-inference"
          },
          "name": "Experimental Design",
          "origin": "catalog",
          "rationale": "The single clearest separator from every neighbouring role."
        }
      ]
    },
    {
      "label": "Primary",
      "layer": "L1",
      "skills": [
        {
          "dimension": {
            "display_name": "Applied ML Modeling \u0026 Methods",
            "slug": "applied-ml-modeling-methods"
          },
          "name": "Machine Learning",
          "origin": "catalog",
          "rationale": "Data Scientists build models constantly \u2014 for explanation and forecasting rather than for a production endpoint."
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "Databricks",
          "origin": "catalog",
          "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine."
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "Apache Spark",
          "origin": "catalog",
          "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine."
        },
        {
          "dimension": {
            "display_name": "Applied ML Modeling \u0026 Methods",
            "slug": "applied-ml-modeling-methods"
          },
          "name": "Natural Language Processing",
          "origin": "catalog",
          "rationale": "Data Scientists build models constantly \u2014 for explanation and forecasting rather than for a production endpoint."
        },
        {
          "dimension": {
            "display_name": "Data Visualisation \u0026 Reporting",
            "slug": "data-visualization-reporting"
          },
          "name": "Microsoft Power BI",
          "origin": "catalog",
          "rationale": "Charts and notebooks are the working surface, not a presentation afterthought."
        },
        {
          "dimension": {
            "display_name": "Data Visualisation \u0026 Reporting",
            "slug": "data-visualization-reporting"
          },
          "name": "Tableau",
          "origin": "catalog",
          "rationale": "Charts and notebooks are the working surface, not a presentation afterthought."
        },
        {
          "dimension": {
            "display_name": "Data Visualisation \u0026 Reporting",
            "slug": "data-visualization-reporting"
          },
          "name": "Data Visualization",
          "origin": "catalog",
          "rationale": "Charts and notebooks are the working surface, not a presentation afterthought."
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "ETL",
          "origin": "catalog",
          "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine."
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "Big Data",
          "origin": "catalog",
          "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine."
        }
      ]
    },
    {
      "label": "Environment",
      "layer": "L2",
      "skills": [
        {
          "dimension": {
            "display_name": "Cloud Platforms",
            "slug": "cloud-platforms"
          },
          "name": "AWS",
          "origin": "catalog",
          "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment."
        },
        {
          "dimension": {
            "display_name": "Cloud Platforms",
            "slug": "cloud-platforms"
          },
          "name": "Microsoft Azure",
          "origin": "catalog",
          "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment."
        },
        {
          "dimension": {
            "display_name": "Foundation Model Platforms \u0026 APIs",
            "slug": "foundation-model-platforms"
          },
          "name": "Generative AI",
          "origin": "catalog",
          "rationale": "LLMs have become a routine analysis tool \u2014 for coding assistance, unstructured text coding and summarisation."
        }
      ]
    },
    {
      "label": "Hygiene",
      "layer": "L3",
      "skills": [
        {
          "dimension": {
            "display_name": "Version Control",
            "slug": "version-control"
          },
          "name": "Git",
          "origin": "catalog",
          "rationale": "Hygiene, and the one engineering habit most worth checking in an otherwise research-leaning candidate."
        }
      ]
    }
  ],
  "unmapped_skills": [
    "AI",
    "data science"
  ]
}
API 2 — extract-details
{}
API 3 — final-role-output
{}

LLM Calls

Every model call made for this run, in pipeline order. Click a card to see the model's response.

Loading…