← Back to history

Pipeline run

9fe0543f-d812-4bb4-80df-10c9891a7e6d

Pipeline LLM cost (USD)
API 1: $0.0037 API 2: $0.0000 API 3: $0.0000 Total: $0.0037

Client output enrichment

v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA description
Nature of work
—
no_db_connection
Tech stack maturity
Mainstream Modern
AI index (0 = no AI use, 5 = totally AI-dependent · v2.1)
0.20 / 5
· Title match
✓ Has AI skill
· AI skill (primary)
· AI skill (secondary)
· On AI team
· Builds AI products
vocab breakdown (legacy)
Assistants (×1): —
Frameworks (×2): —
Models / concepts (×3): AI, Generative AI, Machine Learning
Evidence — skills matched in JD (10)
Python SQL Databricks Apache Spark Microsoft Power BI Tableau ETL AWS Microsoft Azure Git
Skill cluster (0 dimension groups, role-scoped)
No dimension groups computed for this JD.
Status: jd_deconstruction_v4_done Created: 2026-09-03T20:25:30.791660Z Updated: 2026-09-04T23:18:29.922571Z Run duration: 25885 ms
Flow v4 catalog deconstruction skill_library_v4

1 LLM mention extraction + deterministic catalog sweep

2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)

3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim

Role Chosen role & resolution

Data Scientist

in_db body_evidence_advisor family · data-scientist

slug: data-scientist · role_id: 293411ee-d03b-59f7-a4e0-60756ab78387 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)

The role is for a data scientist with a focus on machine learning and data analysis, suitable for candidates with an advanced degree or a bachelor's degree plus two years of relevant experience. The core technology stack includes Python, SQL, cloud platforms like AWS or Azure, and tools such as Databricks and Spark. Distinctive requirements include expertise in machine learning algorithms and experience with data visualization tools like Power BI or Tableau. The position may involve collaboration in cross-functional teams and requires strong problem-solving and communication skills.

Job description

REQUIREMENTS:
• Advanced degree with no prior experience required, or Bachelor's degree plus 2 years of experience in applied data science or related role
• Preferred Fields of Study: Data Science, Computer Science, Statistics, Mathematics, Engineering, or related quantitative field
• Additional relevant certifications in machine learning, AI, or data science a plus
• Strong proficiency in Python, SQL, and data analysis tools
• Expertise in machine learning algorithms, statistical modeling, and data visualization
• Experience with cloud platforms (AWS, Azure) and modern data architectures
• Proficient in data engineering and ETL processes
• Experience with Databricks, Spark, and big data technologies
• Knowledge of generative AI and natural language processing
• Strong understanding of statistical concepts and experimental design
• Experience with data visualization tools (Power BI, Tableau)
• Excellent problem-solving and analytical skills
• Strong communication skills and ability to present complex findings to stakeholders
• Experience working in cross-functional matrix environments
• Demonstrated success in delivering data-driven solutions
• Ability to work independently and collaborate effectively with teams
• Experience with version control systems (e.g., Git)
• Travel may be required (0-25%)

Skills from this JD

L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.

L0 · Anchor 2 skill(s)
Python Programming Languages Python and SQL universally; R in research-leaning and life-sciences teams.
SQL Programming Languages Python and SQL universally; R in research-leaning and life-sciences teams.
L1 · Primary 5 skill(s)
Databricks Data Processing & Pipeline Frameworks pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.
Apache Spark Data Processing & Pipeline Frameworks pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.
Microsoft Power BI Data Visualisation & Reporting Charts and notebooks are the working surface, not a presentation afterthought.
Tableau Data Visualisation & Reporting Charts and notebooks are the working surface, not a presentation afterthought.
ETL Data Processing & Pipeline Frameworks pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.
L2 · Environment 2 skill(s)
AWS Cloud Platforms Notebooks, warehouses and compute increasingly live in a managed environment.
Microsoft Azure Cloud Platforms Notebooks, warehouses and compute increasingly live in a managed environment.
L3 · Hygiene 1 skill(s)
Git Version Control Hygiene, and the one engineering habit most worth checking in an otherwise research-leaning candidate.
Secondary nice-to-have · not a drop-order layer
Machine Learning AI GENERATED not_in_catalog

This skill is in Secondary as the JD says "Additional relevant certifications in machine learning, AI, or data science a plus" - it isn't in this role's skill catalog yet, so it's tracked for review.

Natural Language Processing AI GENERATED not_in_catalog

This skill is in Secondary as the JD says "Knowledge of generative AI and natural language processing" - it isn't in this role's skill catalog yet, so it's tracked for review.

Consider adding canon exemplars the JD missed

A/B Testing — Experimentation & Causal Inference (L0)

R — Programming Languages (L0)

Hypothesis Testing — Statistical Modeling & Inference (L0)

Data Storytelling — Analytical Communication & Decision Support (L1)

Gradient Boosting — Applied ML Modeling & Methods (L1)

PySpark — Data Processing & Pipeline Frameworks (L1)

Unmapped not resolved, not classified — quarantined for review
AIdata sciencegenerative AI

Library artifacts (this run)

No artifact rows for this run.
nano JD Parser — gpt-4.1-nano click to toggle
JD type fail
Show raw JSON
{
  "JD_type": "fail"
}
API 1 — extract-from-jd click to toggle
{
  "catalog": "v4",
  "consider_adding": [
    {
      "dimension": "Experimentation \u0026 Causal Inference",
      "reason": "In the role canon\u0027s Experimentation \u0026 Causal Inference - adding it narrows the candidate pool.",
      "skill": "A/B Testing",
      "tier": "L0"
    },
    {
      "dimension": "Programming Languages",
      "reason": "In the role canon\u0027s Programming Languages - adding it narrows the candidate pool.",
      "skill": "R",
      "tier": "L0"
    },
    {
      "dimension": "Statistical Modeling \u0026 Inference",
      "reason": "In the role canon\u0027s Statistical Modeling \u0026 Inference - adding it narrows the candidate pool.",
      "skill": "Hypothesis Testing",
      "tier": "L0"
    },
    {
      "dimension": "Analytical Communication \u0026 Decision Support",
      "reason": "In the role canon\u0027s Analytical Communication \u0026 Decision Support - adding it narrows the candidate pool.",
      "skill": "Data Storytelling",
      "tier": "L1"
    },
    {
      "dimension": "Applied ML Modeling \u0026 Methods",
      "reason": "In the role canon\u0027s Applied ML Modeling \u0026 Methods - adding it narrows the candidate pool.",
      "skill": "Gradient Boosting",
      "tier": "L1"
    },
    {
      "dimension": "Data Processing \u0026 Pipeline Frameworks",
      "reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
      "skill": "PySpark",
      "tier": "L1"
    }
  ],
  "final_skills": [
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams.",
      "skill_id": "646cb350-7c71-5ed2-971e-d141c566ce6e",
      "skill_name": "Python"
    },
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams.",
      "skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
      "skill_name": "SQL"
    },
    {
      "dimension": {
        "display_name": "Cloud Platforms",
        "slug": "cloud-platforms"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment.",
      "skill_id": "78c5a4b9-2ab5-5279-a688-46f4c3767c3f",
      "skill_name": "AWS"
    },
    {
      "dimension": {
        "display_name": "Cloud Platforms",
        "slug": "cloud-platforms"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment.",
      "skill_id": "060b1d39-d4d8-581d-9d96-65cd45d7cea4",
      "skill_name": "Microsoft Azure"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.",
      "skill_id": "f621b9a9-434a-5aae-a79f-5d858d66973d",
      "skill_name": "Databricks"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.",
      "skill_id": "f885334e-9062-5967-bc62-b5b5a27664cf",
      "skill_name": "Apache Spark"
    },
    {
      "dimension": {
        "display_name": "Data Visualisation \u0026 Reporting",
        "slug": "data-visualization-reporting"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Charts and notebooks are the working surface, not a presentation afterthought.",
      "skill_id": "441f4640-fee5-5dfd-aa1a-b3213839ffef",
      "skill_name": "Microsoft Power BI"
    },
    {
      "dimension": {
        "display_name": "Data Visualisation \u0026 Reporting",
        "slug": "data-visualization-reporting"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Charts and notebooks are the working surface, not a presentation afterthought.",
      "skill_id": "b5b389ea-14af-51b7-a7db-416318e9719b",
      "skill_name": "Tableau"
    },
    {
      "dimension": {
        "display_name": "Version Control",
        "slug": "version-control"
      },
      "is_primary": false,
      "layer": "L3",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Hygiene, and the one engineering habit most worth checking in an otherwise research-leaning candidate.",
      "skill_id": "69a9e960-a4b1-5e7e-b55c-04980465e177",
      "skill_name": "Git"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine.",
      "skill_id": "16a8d754-3339-58e3-978c-7e033bb9b94d",
      "skill_name": "ETL"
    }
  ],
  "history_run_id": null,
  "jd_parameters": {
    "certifications": [],
    "clientDetails": null,
    "company": null,
    "companySize": null,
    "ctc": {
      "currency": null,
      "max": null,
      "min": null,
      "period": null,
      "raw": null
    },
    "educationRequirements": [],
    "experience": {
      "max": null,
      "min": null,
      "raw": null
    },
    "industryDomain": null,
    "knockouts": {
      "certifications": [],
      "ctc": {
        "currency": null,
        "max": null,
        "min": null
      },
      "educationRequirements": [],
      "experience": {
        "max": null,
        "min": null
      },
      "location": [],
      "noticePeriod": null
    },
    "locations": [],
    "noticePeriod": null,
    "openToRelocate": false,
    "role": null,
    "roleSynonyms": []
  },
  "jd_summary": "The role is for a data scientist with a focus on machine learning and data analysis, suitable for candidates with an advanced degree or a bachelor\u0027s degree plus two years of relevant experience. The core technology stack includes Python, SQL, cloud platforms like AWS or Azure, and tools such as Databricks and Spark. Distinctive requirements include expertise in machine learning algorithms and experience with data visualization tools like Power BI or Tableau. The position may involve collaboration in cross-functional teams and requires strong problem-solving and communication skills.",
  "layer_conflicts": [],
  "nano_parsed": {
    "JD_type": "fail"
  },
  "pipeline": "v4",
  "rejected": false,
  "rejection_code": null,
  "rejection_reason": null,
  "role": {
    "canonical_name": "Data Scientist",
    "family": "data-scientist",
    "match_method": "body_evidence_advisor",
    "resolution": "in_db",
    "role_id": "293411ee-d03b-59f7-a4e0-60756ab78387",
    "similarity": null,
    "slug": "data-scientist"
  },
  "role_coverage": [
    {
      "is_jd_role": false,
      "name": "Data Engineer",
      "pct": 0.48,
      "role": "data-engineer",
      "share": 59.1
    },
    {
      "is_jd_role": true,
      "name": "Data Scientist",
      "pct": 0.41,
      "role": "data-scientist",
      "share": 27.5
    },
    {
      "is_jd_role": false,
      "name": "ETL Developer",
      "pct": 0.3,
      "role": "etl-developer",
      "share": 13.4
    }
  ],
  "run_id": "jdv4-185ac22e5878",
  "secondary_meta": [
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"Additional relevant certifications in machine learning, AI, or data science a plus\" - it isn\u0027t in this role\u0027s skill catalog yet, so it\u0027s tracked for review.",
      "jd_quote": "Additional relevant certifications in machine learning, AI, or data science a plus",
      "origin": "ai_predicted",
      "provenance": "listed",
      "reason_code": "not_in_catalog",
      "skill": "Machine Learning"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"Knowledge of generative AI and natural language processing\" - it isn\u0027t in this role\u0027s skill catalog yet, so it\u0027s tracked for review.",
      "jd_quote": "Knowledge of generative AI and natural language processing",
      "origin": "ai_predicted",
      "provenance": "listed",
      "reason_code": "not_in_catalog",
      "skill": "Natural Language Processing"
    }
  ],
  "secondary_skills": [
    "Machine Learning",
    "Natural Language Processing"
  ],
  "skill_layers": [
    {
      "label": "Anchor",
      "layer": "L0",
      "skills": [
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "Python",
          "origin": "catalog",
          "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams."
        },
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "SQL",
          "origin": "catalog",
          "rationale": "Python and SQL universally; R in research-leaning and life-sciences teams."
        }
      ]
    },
    {
      "label": "Primary",
      "layer": "L1",
      "skills": [
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "Databricks",
          "origin": "catalog",
          "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine."
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "Apache Spark",
          "origin": "catalog",
          "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine."
        },
        {
          "dimension": {
            "display_name": "Data Visualisation \u0026 Reporting",
            "slug": "data-visualization-reporting"
          },
          "name": "Microsoft Power BI",
          "origin": "catalog",
          "rationale": "Charts and notebooks are the working surface, not a presentation afterthought."
        },
        {
          "dimension": {
            "display_name": "Data Visualisation \u0026 Reporting",
            "slug": "data-visualization-reporting"
          },
          "name": "Tableau",
          "origin": "catalog",
          "rationale": "Charts and notebooks are the working surface, not a presentation afterthought."
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "ETL",
          "origin": "catalog",
          "rationale": "pandas is the default working surface and Spark the escape hatch when the dataset outgrows one machine."
        }
      ]
    },
    {
      "label": "Environment",
      "layer": "L2",
      "skills": [
        {
          "dimension": {
            "display_name": "Cloud Platforms",
            "slug": "cloud-platforms"
          },
          "name": "AWS",
          "origin": "catalog",
          "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment."
        },
        {
          "dimension": {
            "display_name": "Cloud Platforms",
            "slug": "cloud-platforms"
          },
          "name": "Microsoft Azure",
          "origin": "catalog",
          "rationale": "Notebooks, warehouses and compute increasingly live in a managed environment."
        }
      ]
    },
    {
      "label": "Hygiene",
      "layer": "L3",
      "skills": [
        {
          "dimension": {
            "display_name": "Version Control",
            "slug": "version-control"
          },
          "name": "Git",
          "origin": "catalog",
          "rationale": "Hygiene, and the one engineering habit most worth checking in an otherwise research-leaning candidate."
        }
      ]
    }
  ],
  "unmapped_skills": [
    "AI",
    "data science",
    "generative AI"
  ]
}
API 2 — extract-details
{}
API 3 — final-role-output
{}

LLM Calls

Every model call made for this run, in pipeline order. Click a card to see the model's response.

Loading…