<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://donmangudata-ops.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://donmangudata-ops.github.io/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-09-29T05:54:56+00:00</updated><id>https://donmangudata-ops.github.io/feed.xml</id><title type="html">Don Mangu</title><subtitle>Practical guides on job board APIs, hiring signals and website data in Python, with the free endpoints first and hosted Apify Actors where they save time.</subtitle><author><name>Don Mangu</name></author><entry><title type="html">Ashby Job Board API: Read Open Jobs and Salary Ranges in Python</title><link href="https://donmangudata-ops.github.io/ashby-job-board-api-python/" rel="alternate" type="text/html" title="Ashby Job Board API: Read Open Jobs and Salary Ranges in Python" /><published>2026-09-29T00:00:00+00:00</published><updated>2026-09-29T00:00:00+00:00</updated><id>https://donmangudata-ops.github.io/ashby-job-board-api-python</id><content type="html" xml:base="https://donmangudata-ops.github.io/ashby-job-board-api-python/"><![CDATA[<p><em>Disclosure: I built the Apify Actors mentioned in the second half, and they are paid. This article was drafted with an AI assistant. The endpoint was checked against live Ashby boards on September 26, 2026. Actor prices and field names were checked on September 27, 2026.</em></p>

<p>Ashby has become a common applicant tracking system at startups and AI companies. Its hosted job boards live at <code class="language-plaintext highlighter-rouge">jobs.ashbyhq.com/{board}</code>, and Ashby documents a public posting API that returns the same jobs as JSON. For job data it also has something the older systems often lack, structured salary ranges, when the company chooses to publish them.</p>

<h2 id="the-endpoint">The endpoint</h2>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>GET https://api.ashbyhq.com/posting-api/job-board/{board}?includeCompensation=true
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">{board}</code> is the part after <code class="language-plaintext highlighter-rouge">jobs.ashbyhq.com/</code>. For <code class="language-plaintext highlighter-rouge">jobs.ashbyhq.com/openai</code> it is <code class="language-plaintext highlighter-rouge">openai</code>. No key is needed to read published jobs. An unknown board returns HTTP 404.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">requests</span>

<span class="k">def</span> <span class="nf">ashby_jobs</span><span class="p">(</span><span class="n">board</span><span class="p">):</span>
    <span class="n">url</span> <span class="o">=</span> <span class="sa">f</span><span class="s">"https://api.ashbyhq.com/posting-api/job-board/</span><span class="si">{</span><span class="n">board</span><span class="si">}</span><span class="s">"</span>
    <span class="n">resp</span> <span class="o">=</span> <span class="n">requests</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="n">url</span><span class="p">,</span> <span class="n">params</span><span class="o">=</span><span class="p">{</span><span class="s">"includeCompensation"</span><span class="p">:</span> <span class="s">"true"</span><span class="p">},</span> <span class="n">timeout</span><span class="o">=</span><span class="mi">30</span><span class="p">)</span>
    <span class="k">if</span> <span class="n">resp</span><span class="p">.</span><span class="n">status_code</span> <span class="o">==</span> <span class="mi">404</span><span class="p">:</span>
        <span class="k">return</span> <span class="bp">None</span>
    <span class="n">resp</span><span class="p">.</span><span class="n">raise_for_status</span><span class="p">()</span>
    <span class="k">return</span> <span class="n">resp</span><span class="p">.</span><span class="n">json</span><span class="p">()[</span><span class="s">"jobs"</span><span class="p">]</span>

<span class="n">jobs</span> <span class="o">=</span> <span class="n">ashby_jobs</span><span class="p">(</span><span class="s">"ashby"</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="nb">len</span><span class="p">(</span><span class="n">jobs</span><span class="p">),</span> <span class="s">"jobs"</span><span class="p">)</span>
<span class="k">for</span> <span class="n">j</span> <span class="ow">in</span> <span class="n">jobs</span><span class="p">[:</span><span class="mi">5</span><span class="p">]:</span>
    <span class="k">print</span><span class="p">(</span><span class="n">j</span><span class="p">[</span><span class="s">"title"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">j</span><span class="p">[</span><span class="s">"department"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">j</span><span class="p">[</span><span class="s">"location"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">j</span><span class="p">[</span><span class="s">"workplaceType"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">j</span><span class="p">[</span><span class="s">"publishedAt"</span><span class="p">][:</span><span class="mi">10</span><span class="p">])</span>
</code></pre></div></div>

<p>The whole board comes back in one response. OpenAI’s board had 830 jobs when I checked, with no paging.</p>

<h2 id="fields-worth-knowing">Fields worth knowing</h2>

<ul>
  <li><code class="language-plaintext highlighter-rouge">title</code>, <code class="language-plaintext highlighter-rouge">department</code>, <code class="language-plaintext highlighter-rouge">team</code>, <code class="language-plaintext highlighter-rouge">employmentType</code></li>
  <li><code class="language-plaintext highlighter-rouge">location</code> plus <code class="language-plaintext highlighter-rouge">secondaryLocations</code> for jobs open in several places</li>
  <li><code class="language-plaintext highlighter-rouge">isRemote</code> and <code class="language-plaintext highlighter-rouge">workplaceType</code></li>
  <li><code class="language-plaintext highlighter-rouge">publishedAt</code>, an ISO timestamp</li>
  <li><code class="language-plaintext highlighter-rouge">jobUrl</code> and <code class="language-plaintext highlighter-rouge">applyUrl</code></li>
  <li><code class="language-plaintext highlighter-rouge">descriptionPlain</code> and <code class="language-plaintext highlighter-rouge">descriptionHtml</code></li>
  <li><code class="language-plaintext highlighter-rouge">isListed</code>, which is false for unlisted jobs that are reachable only by link</li>
  <li><code class="language-plaintext highlighter-rouge">compensation</code>, only when you pass <code class="language-plaintext highlighter-rouge">includeCompensation=true</code></li>
</ul>

<h3 id="reading-the-salary">Reading the salary</h3>

<p><code class="language-plaintext highlighter-rouge">compensation.summaryComponents</code> is a list of parts: salary, bonus, equity and so on. Each has <code class="language-plaintext highlighter-rouge">compensationType</code>, <code class="language-plaintext highlighter-rouge">interval</code>, <code class="language-plaintext highlighter-rouge">currencyCode</code>, <code class="language-plaintext highlighter-rouge">minValue</code> and <code class="language-plaintext highlighter-rouge">maxValue</code>. The salary is the part whose type is <code class="language-plaintext highlighter-rouge">Salary</code>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">salary_range</span><span class="p">(</span><span class="n">job</span><span class="p">):</span>
    <span class="n">comp</span> <span class="o">=</span> <span class="n">job</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"compensation"</span><span class="p">)</span> <span class="ow">or</span> <span class="p">{}</span>
    <span class="k">for</span> <span class="n">part</span> <span class="ow">in</span> <span class="n">comp</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"summaryComponents"</span><span class="p">)</span> <span class="ow">or</span> <span class="p">[]:</span>
        <span class="k">if</span> <span class="n">part</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"compensationType"</span><span class="p">)</span> <span class="o">==</span> <span class="s">"Salary"</span> <span class="ow">and</span> <span class="n">part</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"minValue"</span><span class="p">)</span> <span class="ow">is</span> <span class="ow">not</span> <span class="bp">None</span><span class="p">:</span>
            <span class="k">return</span> <span class="n">part</span><span class="p">[</span><span class="s">"minValue"</span><span class="p">],</span> <span class="n">part</span><span class="p">[</span><span class="s">"maxValue"</span><span class="p">],</span> <span class="n">part</span><span class="p">[</span><span class="s">"currencyCode"</span><span class="p">],</span> <span class="n">part</span><span class="p">[</span><span class="s">"interval"</span><span class="p">]</span>
    <span class="k">return</span> <span class="bp">None</span>

<span class="k">for</span> <span class="n">j</span> <span class="ow">in</span> <span class="n">jobs</span><span class="p">:</span>
    <span class="n">s</span> <span class="o">=</span> <span class="n">salary_range</span><span class="p">(</span><span class="n">j</span><span class="p">)</span>
    <span class="k">if</span> <span class="n">s</span><span class="p">:</span>
        <span class="k">print</span><span class="p">(</span><span class="n">j</span><span class="p">[</span><span class="s">"title"</span><span class="p">],</span> <span class="n">s</span><span class="p">)</span>
</code></pre></div></div>

<p>On Ashby’s own board, an engineering manager role in the EU returned <code class="language-plaintext highlighter-rouge">(110000, 185000, 'EUR', '1 YEAR')</code>. There is also <code class="language-plaintext highlighter-rouge">compensationTierSummary</code>, a display string such as “€110K - €185K • Offers Equity”, which is handy for a UI but not for math.</p>

<p>A useful side effect: <code class="language-plaintext highlighter-rouge">publishedAt</code> shows how long a job has been open. The same board had a job first published in March 2024 and still listed in September 2026. Old open jobs are normal on career pages, and job aggregators often drop them.</p>

<h2 id="where-this-stops-being-enough">Where this stops being enough</h2>

<p>For one board this is ten lines of code. The trouble starts with a list:</p>

<ul>
  <li><strong>Board names.</strong> A list of company websites does not tell you the Ashby board name, or whether the company uses Ashby at all.</li>
  <li><strong>Mixed systems.</strong> Few lists are all Ashby. Greenhouse and Lever are common in the same segment, and European companies often use Personio, Teamtailor or Recruitee.</li>
  <li><strong>Other salary formats.</strong> Ashby gives structured pay. Greenhouse usually does not in its list endpoint, and Lever only sometimes. Comparing pay across companies needs one format.</li>
  <li><strong>What changed.</strong> New, closed and reposted jobs need stored history.</li>
</ul>

<h2 id="a-hosted-option-for-many-ashby-boards">A hosted option for many Ashby boards</h2>

<p><a href="https://apify.com/conserving_celerytop/ashby-jobs-api">Ashby Jobs API</a> is an Apify Actor I built for this. You give it Ashby board links, company websites or plain names, it reads the same public posting API for each one, and it returns every job in a fixed format. Ashby salaries arrive as <code class="language-plaintext highlighter-rouge">salaryMin</code>, <code class="language-plaintext highlighter-rouge">salaryMax</code>, <code class="language-plaintext highlighter-rouge">salaryCurrency</code> and <code class="language-plaintext highlighter-rouge">salaryPeriod</code>, with <code class="language-plaintext highlighter-rouge">salaryRanges</code> holding one range per pay tier, plus <code class="language-plaintext highlighter-rouge">salaryAnnualMin</code> and <code class="language-plaintext highlighter-rouge">salaryAnnualMax</code> converted to a yearly figure.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"apify-client&gt;=3.2,&lt;4"</span>
<span class="nb">export </span><span class="nv">APIFY_TOKEN</span><span class="o">=</span>&lt;YOUR_APIFY_TOKEN&gt;
</code></pre></div></div>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">os</span>
<span class="kn">from</span> <span class="nn">decimal</span> <span class="kn">import</span> <span class="n">Decimal</span>
<span class="kn">from</span> <span class="nn">apify_client</span> <span class="kn">import</span> <span class="n">ApifyClient</span>

<span class="n">client</span> <span class="o">=</span> <span class="n">ApifyClient</span><span class="p">(</span><span class="n">os</span><span class="p">.</span><span class="n">environ</span><span class="p">[</span><span class="s">"APIFY_TOKEN"</span><span class="p">])</span>

<span class="n">run</span> <span class="o">=</span> <span class="n">client</span><span class="p">.</span><span class="n">actor</span><span class="p">(</span><span class="s">"conserving_celerytop/ashby-jobs-api"</span><span class="p">).</span><span class="n">call</span><span class="p">(</span>
    <span class="n">run_input</span><span class="o">=</span><span class="p">{</span>
        <span class="s">"companies"</span><span class="p">:</span> <span class="p">[</span>
            <span class="s">"https://jobs.ashbyhq.com/openai"</span><span class="p">,</span>
            <span class="s">"https://jobs.ashbyhq.com/ashby"</span><span class="p">,</span>
        <span class="p">],</span>
        <span class="s">"hasSalary"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
        <span class="s">"minAnnualSalary"</span><span class="p">:</span> <span class="mi">150000</span><span class="p">,</span>
        <span class="s">"minAnnualSalaryCurrency"</span><span class="p">:</span> <span class="s">"USD"</span><span class="p">,</span>
    <span class="p">},</span>
    <span class="n">max_total_charge_usd</span><span class="o">=</span><span class="n">Decimal</span><span class="p">(</span><span class="s">"0.50"</span><span class="p">),</span>
<span class="p">)</span>

<span class="k">for</span> <span class="n">r</span> <span class="ow">in</span> <span class="n">client</span><span class="p">.</span><span class="n">dataset</span><span class="p">(</span><span class="n">run</span><span class="p">.</span><span class="n">default_dataset_id</span><span class="p">).</span><span class="n">iterate_items</span><span class="p">():</span>
    <span class="k">if</span> <span class="n">r</span><span class="p">[</span><span class="s">"rowType"</span><span class="p">]</span> <span class="o">==</span> <span class="s">"job"</span><span class="p">:</span>
        <span class="k">print</span><span class="p">(</span><span class="n">r</span><span class="p">[</span><span class="s">"companySlug"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">r</span><span class="p">[</span><span class="s">"title"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">r</span><span class="p">[</span><span class="s">"salaryAnnualMin"</span><span class="p">],</span> <span class="s">"-"</span><span class="p">,</span> <span class="n">r</span><span class="p">[</span><span class="s">"salaryAnnualMax"</span><span class="p">],</span> <span class="n">r</span><span class="p">[</span><span class="s">"salaryCurrency"</span><span class="p">])</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">hasSalary</code> keeps only jobs with published pay, from the board field or written in the job text. <code class="language-plaintext highlighter-rouge">minAnnualSalary</code> compares the yearly figure in the currency you name. There are no exchange rates, so jobs paid in euros or pounds are left out of a USD filter, and the <code class="language-plaintext highlighter-rouge">warning</code> field counts them. Set the currency to match the market you are looking at.</p>

<p>To watch a set of Ashby companies, add <code class="language-plaintext highlighter-rouge">"onlyNewJobs": true</code> and <code class="language-plaintext highlighter-rouge">"monitorName": "ashby-watch"</code>, save the input as a task, and schedule it. Later runs return only new and closed jobs.</p>

<p>If your list mixes Ashby with Greenhouse, Lever or Workday, <a href="https://apify.com/conserving_celerytop/live-career-page-jobs-api">ATS Jobs API</a> from the same author reads 22 systems with the same fields, at $0.045 per company.</p>

<h2 id="pricing">Pricing</h2>

<p>Free plan prices on the Apify Store on September 27, 2026. Check the Store page before relying on them.</p>

<ul>
  <li>$0.10 per company, including up to 1,000 of its open jobs ($0.09 on Scale, $0.07 on Business). Each further 1,000 jobs of that company is $0.045.</li>
  <li>A later check with <code class="language-plaintext highlighter-rouge">onlyNewJobs</code> is $0.002 per 1,000 open jobs on the board.</li>
  <li>Descriptions cost nothing extra.</li>
  <li>A name that turns out not to be on Ashby is still charged, because the board was looked up. Links to other job boards, websites that cannot be read, duplicates and invalid entries are free.</li>
  <li>Paid Apify plans pay less per company. The free plan gives $5 of credit a month, which covers about 50 companies.</li>
</ul>

<p>The example above, two companies, costs $0.20.</p>

<h2 id="when-not-to-use-it">When not to use it</h2>

<ul>
  <li>One or two Ashby boards: the public endpoint is free and simple.</li>
  <li>A keyword search across every Ashby company: this reads the companies you list. It is not a search engine over all boards.</li>
  <li>Your own Ashby hiring data (candidates, interviews): that is Ashby’s authenticated API, with a key from your Ashby admin.</li>
</ul>

<h2 id="faq">FAQ</h2>

<p><strong>Does Ashby’s job posting API need a key?</strong>
No, not for published jobs.</p>

<p><strong>Why is <code class="language-plaintext highlighter-rouge">compensation</code> missing?</strong>
Either you did not pass <code class="language-plaintext highlighter-rouge">includeCompensation=true</code>, or the company does not show pay on that job.</p>

<p><strong>Are unlisted jobs included?</strong>
The API can return jobs with <code class="language-plaintext highlighter-rouge">isListed</code> set to false. Filter on it if you only want jobs shown on the public board. The Actor leaves unlisted jobs out.</p>

<hr />

<p><em>Ashby is a trademark of its owner. This article and the Actors are not affiliated with or endorsed by Ashby. The Actors read only jobs published on public boards, with no login.</em></p>]]></content><author><name>Don Mangu</name></author><category term="python" /><category term="api" /><category term="web-scraping" /><category term="apify" /><category term="recruitment" /><summary type="html"><![CDATA[Use Ashby's public job posting API to pull open jobs and salary ranges from any Ashby board in Python, then scale it to many companies with one schema.]]></summary></entry><entry><title type="html">Hiring Signals for Sales: Score Your Account List From Job Postings</title><link href="https://donmangudata-ops.github.io/hiring-signals-for-sales/" rel="alternate" type="text/html" title="Hiring Signals for Sales: Score Your Account List From Job Postings" /><published>2026-09-29T00:00:00+00:00</published><updated>2026-09-29T00:00:00+00:00</updated><id>https://donmangudata-ops.github.io/hiring-signals-for-sales</id><content type="html" xml:base="https://donmangudata-ops.github.io/hiring-signals-for-sales/"><![CDATA[<p><em>Disclosure: I built the Apify Actor used in this guide, and it is paid. This article was drafted with an AI assistant. Field names and prices were checked on September 27, 2026.</em></p>

<p>A company that opens three sales roles this month is about to change how it sells. One that posts its first data engineer is about to buy data tools. Job postings are one of the few buying signals a company publishes on purpose, with dates, in public.</p>

<p>Most signal tools sell that data inside a larger platform. If you already have an account list, you can build the core of it yourself: read each account’s careers page, count what they are hiring for, and flag the changes. This guide does that in Python.</p>

<h2 id="which-hiring-signals-matter">Which hiring signals matter</h2>

<p>Pick signals that map to what you sell. Some that hold up in practice:</p>

<ul>
  <li><strong>Hiring in your buyer’s team.</strong> Selling to RevOps? Watch for sales and RevOps roles. Selling dev tools? Watch engineering.</li>
  <li><strong>A first hire in a function.</strong> “Founding data engineer” or “first marketing hire” means someone is about to pick tools with no incumbent.</li>
  <li><strong>New leadership.</strong> An open VP or director role often comes before a budget and a vendor review.</li>
  <li><strong>Speed.</strong> Many roles posted in the last 30 days, compared with the total open.</li>
  <li><strong>Expansion.</strong> Jobs in a country where the account had none before.</li>
  <li><strong>Tools in the job text.</strong> A post that asks for Salesforce, Snowflake or HubSpot experience tells you what they already run.</li>
</ul>

<p>None of these proves intent. They tell you where to look first, and they give the first line of an email something true to say.</p>

<h2 id="step-1-one-summary-row-per-account">Step 1: one summary row per account</h2>

<p><a href="https://apify.com/conserving_celerytop/live-career-page-jobs-api">ATS Jobs API</a> reads the public careers pages of the companies you give it: Greenhouse, Lever, Ashby, Workday and 18 other job board systems. With <code class="language-plaintext highlighter-rouge">outputMode</code> set to <code class="language-plaintext highlighter-rouge">companies</code>, it returns one summary row per account instead of one row per job.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"apify-client&gt;=3.2,&lt;4"</span>
<span class="nb">export </span><span class="nv">APIFY_TOKEN</span><span class="o">=</span>&lt;YOUR_APIFY_TOKEN&gt;
</code></pre></div></div>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">os</span>
<span class="kn">from</span> <span class="nn">decimal</span> <span class="kn">import</span> <span class="n">Decimal</span>
<span class="kn">from</span> <span class="nn">apify_client</span> <span class="kn">import</span> <span class="n">ApifyClient</span>

<span class="n">client</span> <span class="o">=</span> <span class="n">ApifyClient</span><span class="p">(</span><span class="n">os</span><span class="p">.</span><span class="n">environ</span><span class="p">[</span><span class="s">"APIFY_TOKEN"</span><span class="p">])</span>

<span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="s">"accounts.txt"</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>          <span class="c1"># one website or domain per line
</span>    <span class="n">accounts</span> <span class="o">=</span> <span class="p">[</span><span class="n">line</span><span class="p">.</span><span class="n">strip</span><span class="p">()</span> <span class="k">for</span> <span class="n">line</span> <span class="ow">in</span> <span class="n">f</span> <span class="k">if</span> <span class="n">line</span><span class="p">.</span><span class="n">strip</span><span class="p">()]</span>

<span class="n">run</span> <span class="o">=</span> <span class="n">client</span><span class="p">.</span><span class="n">actor</span><span class="p">(</span><span class="s">"conserving_celerytop/live-career-page-jobs-api"</span><span class="p">).</span><span class="n">call</span><span class="p">(</span>
    <span class="n">run_input</span><span class="o">=</span><span class="p">{</span>
        <span class="s">"companies"</span><span class="p">:</span> <span class="n">accounts</span><span class="p">[:</span><span class="mi">500</span><span class="p">],</span>
        <span class="s">"outputMode"</span><span class="p">:</span> <span class="s">"companies"</span><span class="p">,</span>
        <span class="s">"includeDescription"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>     <span class="c1"># adds topTools and first hires found in the job text
</span>    <span class="p">},</span>
    <span class="n">max_total_charge_usd</span><span class="o">=</span><span class="n">Decimal</span><span class="p">(</span><span class="s">"25.00"</span><span class="p">),</span>  <span class="c1"># a hard cap, not the price; see "What it costs"
</span><span class="p">)</span>

<span class="n">rows</span> <span class="o">=</span> <span class="p">[</span><span class="n">r</span> <span class="k">for</span> <span class="n">r</span> <span class="ow">in</span> <span class="n">client</span><span class="p">.</span><span class="n">dataset</span><span class="p">(</span><span class="n">run</span><span class="p">.</span><span class="n">default_dataset_id</span><span class="p">).</span><span class="n">iterate_items</span><span class="p">()</span> <span class="k">if</span> <span class="n">r</span><span class="p">[</span><span class="s">"rowType"</span><span class="p">]</span> <span class="o">==</span> <span class="s">"company"</span><span class="p">]</span>
</code></pre></div></div>

<p>Each summary row includes:</p>

<table>
  <thead>
    <tr>
      <th>Field</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">openJobs</code></td>
      <td>Open roles</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">jobsPostedLast7Days</code>, <code class="language-plaintext highlighter-rouge">jobsPostedLast30Days</code></td>
      <td>Recent roles (null if the board gives no dates)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">functionCounts</code>, <code class="language-plaintext highlighter-rouge">functionCountsLast30Days</code></td>
      <td>Roles per function, such as <code class="language-plaintext highlighter-rouge">{"sales": 12, "engineering": 40}</code></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">salesShare</code>, <code class="language-plaintext highlighter-rouge">engineeringShare</code></td>
      <td>Share of sales and engineering roles</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">leadershipRoles</code></td>
      <td>Up to 5 open director, VP and C-level roles, newest first</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">firstHireRoles</code></td>
      <td>Up to 5 roles described as a first or founding hire</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">topTools</code></td>
      <td>Tools named most often in the job text, with a category</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">countries</code></td>
      <td>Where the open roles are</td>
    </tr>
  </tbody>
</table>

<p>Accounts it cannot resolve come back as rows with <code class="language-plaintext highlighter-rouge">rowType</code> set to <code class="language-plaintext highlighter-rouge">status</code> and a reason in <code class="language-plaintext highlighter-rouge">companyStatus</code>, such as <code class="language-plaintext highlighter-rouge">no_job_board_found</code>. Websites work when the site links to its job board; a board link always works best.</p>

<h2 id="step-2-score-the-accounts">Step 2: score the accounts</h2>

<p>Here is a simple score for a team that sells to sales leaders. Change the weights to fit your product.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">score</span><span class="p">(</span><span class="n">r</span><span class="p">,</span> <span class="n">function</span><span class="o">=</span><span class="s">"sales"</span><span class="p">):</span>
    <span class="n">recent</span> <span class="o">=</span> <span class="p">(</span><span class="n">r</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"functionCountsLast30Days"</span><span class="p">)</span> <span class="ow">or</span> <span class="p">{}).</span><span class="n">get</span><span class="p">(</span><span class="n">function</span><span class="p">,</span> <span class="mi">0</span><span class="p">)</span>
    <span class="n">total</span> <span class="o">=</span> <span class="p">(</span><span class="n">r</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"functionCounts"</span><span class="p">)</span> <span class="ow">or</span> <span class="p">{}).</span><span class="n">get</span><span class="p">(</span><span class="n">function</span><span class="p">,</span> <span class="mi">0</span><span class="p">)</span>
    <span class="n">leaders</span> <span class="o">=</span> <span class="p">[</span><span class="n">x</span> <span class="k">for</span> <span class="n">x</span> <span class="ow">in</span> <span class="p">(</span><span class="n">r</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"leadershipRoles"</span><span class="p">)</span> <span class="ow">or</span> <span class="p">[])</span> <span class="k">if</span> <span class="n">x</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"jobFunction"</span><span class="p">)</span> <span class="o">==</span> <span class="n">function</span><span class="p">]</span>
    <span class="n">firsts</span> <span class="o">=</span> <span class="p">[</span><span class="n">x</span> <span class="k">for</span> <span class="n">x</span> <span class="ow">in</span> <span class="p">(</span><span class="n">r</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"firstHireRoles"</span><span class="p">)</span> <span class="ow">or</span> <span class="p">[])</span> <span class="k">if</span> <span class="n">x</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"jobFunction"</span><span class="p">)</span> <span class="o">==</span> <span class="n">function</span><span class="p">]</span>
    <span class="k">return</span> <span class="n">recent</span> <span class="o">*</span> <span class="mi">3</span> <span class="o">+</span> <span class="n">total</span> <span class="o">+</span> <span class="nb">len</span><span class="p">(</span><span class="n">leaders</span><span class="p">)</span> <span class="o">*</span> <span class="mi">10</span> <span class="o">+</span> <span class="nb">len</span><span class="p">(</span><span class="n">firsts</span><span class="p">)</span> <span class="o">*</span> <span class="mi">15</span>

<span class="n">ranked</span> <span class="o">=</span> <span class="nb">sorted</span><span class="p">(</span><span class="n">rows</span><span class="p">,</span> <span class="n">key</span><span class="o">=</span><span class="n">score</span><span class="p">,</span> <span class="n">reverse</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
<span class="k">for</span> <span class="n">r</span> <span class="ow">in</span> <span class="n">ranked</span><span class="p">[:</span><span class="mi">25</span><span class="p">]:</span>
    <span class="n">leads</span> <span class="o">=</span> <span class="s">", "</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="n">x</span><span class="p">[</span><span class="s">"title"</span><span class="p">]</span> <span class="k">for</span> <span class="n">x</span> <span class="ow">in</span> <span class="p">(</span><span class="n">r</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"leadershipRoles"</span><span class="p">)</span> <span class="ow">or</span> <span class="p">[])[:</span><span class="mi">2</span><span class="p">])</span>
    <span class="n">name</span> <span class="o">=</span> <span class="n">r</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"companySlug"</span><span class="p">)</span> <span class="ow">or</span> <span class="n">r</span><span class="p">[</span><span class="s">"company"</span><span class="p">]</span>
    <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'</span><span class="si">{</span><span class="n">name</span><span class="si">:</span><span class="o">&lt;</span><span class="mi">20</span><span class="si">}</span><span class="s"> score=</span><span class="si">{</span><span class="n">score</span><span class="p">(</span><span class="n">r</span><span class="p">)</span><span class="si">:</span><span class="o">&gt;</span><span class="mi">3</span><span class="si">}</span><span class="s">  sales_open=</span><span class="si">{</span><span class="p">(</span><span class="n">r</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"functionCounts"</span><span class="p">)</span> <span class="ow">or</span> <span class="si">{}</span><span class="p">).</span><span class="n">get</span><span class="p">(</span><span class="s">"sales"</span><span class="p">,</span> <span class="mi">0</span><span class="p">)</span><span class="si">:</span><span class="o">&gt;</span><span class="mi">3</span><span class="si">}</span><span class="s">  </span><span class="si">{</span><span class="n">leads</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
</code></pre></div></div>

<p>Push the top 25 to your CRM with the job titles as context. “Saw you’re hiring a VP Sales and four AEs in Austin” beats a generic opener, and it is true.</p>

<h2 id="step-3-a-weekly-list-of-what-changed">Step 3: a weekly list of what changed</h2>

<p>A snapshot is useful once. The value is in change. Add three fields and schedule it:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">run_input</span> <span class="o">=</span> <span class="p">{</span>
    <span class="s">"companies"</span><span class="p">:</span> <span class="n">accounts</span><span class="p">[:</span><span class="mi">500</span><span class="p">],</span>
    <span class="s">"outputMode"</span><span class="p">:</span> <span class="s">"companies"</span><span class="p">,</span>
    <span class="s">"onlyNewJobs"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
    <span class="s">"monitorName"</span><span class="p">:</span> <span class="s">"accounts-weekly"</span><span class="p">,</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The first run records every open job. Each later run adds <code class="language-plaintext highlighter-rouge">newJobs</code>, <code class="language-plaintext highlighter-rouge">closedJobs</code>, <code class="language-plaintext highlighter-rouge">newFunctions</code> (functions with open roles now that had none last time, for example <code class="language-plaintext highlighter-rouge">["sales"]</code>) and <code class="language-plaintext highlighter-rouge">newCountries</code>. Save the input as a task in Apify Console, add a weekly schedule, and send the result to Slack, Google Sheets or a webhook. An account whose <code class="language-plaintext highlighter-rouge">newFunctions</code> contains your buyer’s team is worth a look that week.</p>

<h2 id="what-it-costs">What it costs</h2>

<p>On September 27, 2026 the Store price was $0.045 per account, with up to 1,000 jobs included, and $0.0428, $0.0405 or $0.036 on paid Apify plans. Prices can change, so check the <a href="https://apify.com/conserving_celerytop/live-career-page-jobs-api">Store page</a> before a large run. How it adds up:</p>

<ul>
  <li>The first run costs one company lookup per account. In <code class="language-plaintext highlighter-rouge">companies</code> mode that is the whole price per account, however many jobs it has. So the first run for 500 accounts costs $22.50 on the free plan.</li>
  <li>Job descriptions, needed for <code class="language-plaintext highlighter-rouge">topTools</code> and first hires found in the text, are free on most systems. Workday, Eightfold and a few smaller systems, such as JazzHR and Paylocity, cost $0.01 per started block of 200 jobs described.</li>
  <li>Each weekly check after the first costs $0.002 per 1,000 open jobs on each board (checked September 27, 2026). Most accounts have fewer than 1,000 open jobs, so 500 accounts cost about $1 a week.</li>
  <li><code class="language-plaintext highlighter-rouge">max_total_charge_usd</code> in the code is a hard cap. The run stops before it spends more.</li>
</ul>

<p>Apify’s free plan includes $5 of credit a month.</p>

<h2 id="limits-you-should-know">Limits you should know</h2>

<ul>
  <li><strong>Function labels are read from titles.</strong> Some jobs land in the wrong function, so read the titles before you act on a count. <code class="language-plaintext highlighter-rouge">salesShare</code> and <code class="language-plaintext highlighter-rouge">engineeringShare</code> use simpler keyword rules.</li>
  <li><strong>It reads the companies you list.</strong> It will not find new accounts that are hiring. Bring a list from your CRM or a data provider.</li>
  <li><strong>Not every company has a careers page system.</strong> Small companies that post only on LinkedIn come back as not found.</li>
  <li><strong>Old postings are included.</strong> Some companies keep evergreen roles open all year. <code class="language-plaintext highlighter-rouge">jobsPostedLast30Days</code> separates new hiring from old listings.</li>
</ul>

<h2 id="faq">FAQ</h2>

<p><strong>Is hiring really a buying signal?</strong>
It is a timing signal. It shows where budget and headcount are going, which is often where new tools are bought. Use it to prioritize, not as proof.</p>

<p><strong>Can I get the hiring manager’s name?</strong>
No. The Actor reads job postings only and returns no personal data.</p>

<p><strong>How is this different from a signals platform?</strong>
You get the raw counts and job titles for your own list, at a per-account price, and you decide the scoring. Platforms bundle more sources and contact data.</p>

<hr />

<p><em>This article and the Actor are not affiliated with or endorsed by Greenhouse, Lever, Ashby, Workday or any other job board. The Actor reads only jobs published on public careers pages, with no login.</em></p>]]></content><author><name>Don Mangu</name></author><category term="sales" /><category term="lead-generation" /><category term="python" /><category term="automation" /><category term="apify" /><summary type="html"><![CDATA[Turn job postings into hiring signals for outbound: pull open roles for each account, score them in Python, and get a weekly list of accounts that changed.]]></summary></entry><entry><title type="html">Lever API Job Postings: Pull Open Jobs From Lever With Python</title><link href="https://donmangudata-ops.github.io/lever-postings-api-python/" rel="alternate" type="text/html" title="Lever API Job Postings: Pull Open Jobs From Lever With Python" /><published>2026-09-28T09:00:00+00:00</published><updated>2026-09-28T09:00:00+00:00</updated><id>https://donmangudata-ops.github.io/lever-postings-api-python</id><content type="html" xml:base="https://donmangudata-ops.github.io/lever-postings-api-python/"><![CDATA[<p><em>Disclosure: I built the Apify Actors mentioned in the second half, and they are paid. This article was drafted with an AI assistant. Endpoints were checked against live Lever boards on September 26, 2026. Actor prices and field names were checked on September 27, 2026.</em></p>

<p>“Lever API” means two different things, and search results mix them up. One is the Lever API for customers, which reads candidates, opportunities and interviews and needs an API key from the company’s Lever account. The other is the <strong>Postings API</strong>, which serves the jobs a company publishes on its Lever job site. The Postings API is public for reading, and it is the one you want for job data.</p>

<h2 id="the-endpoint">The endpoint</h2>

<p>Every Lever job site at <code class="language-plaintext highlighter-rouge">jobs.lever.co/{site}</code> has a matching JSON endpoint:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>GET https://api.lever.co/v0/postings/{site}?mode=json
</code></pre></div></div>

<p>Boards hosted in Lever’s EU region live at <code class="language-plaintext highlighter-rouge">jobs.eu.lever.co/{site}</code>, and their API host is <code class="language-plaintext highlighter-rouge">api.eu.lever.co</code>. If a site returns <code class="language-plaintext highlighter-rouge">{"ok": false, "error": "Document not found"}</code> on one host, try the other before you give up on it.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">requests</span>
<span class="kn">from</span> <span class="nn">datetime</span> <span class="kn">import</span> <span class="n">datetime</span><span class="p">,</span> <span class="n">timezone</span>

<span class="k">def</span> <span class="nf">lever_postings</span><span class="p">(</span><span class="n">site</span><span class="p">,</span> <span class="n">eu</span><span class="o">=</span><span class="bp">False</span><span class="p">):</span>
    <span class="n">host</span> <span class="o">=</span> <span class="s">"api.eu.lever.co"</span> <span class="k">if</span> <span class="n">eu</span> <span class="k">else</span> <span class="s">"api.lever.co"</span>
    <span class="n">resp</span> <span class="o">=</span> <span class="n">requests</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="sa">f</span><span class="s">"https://</span><span class="si">{</span><span class="n">host</span><span class="si">}</span><span class="s">/v0/postings/</span><span class="si">{</span><span class="n">site</span><span class="si">}</span><span class="s">"</span><span class="p">,</span> <span class="n">params</span><span class="o">=</span><span class="p">{</span><span class="s">"mode"</span><span class="p">:</span> <span class="s">"json"</span><span class="p">},</span> <span class="n">timeout</span><span class="o">=</span><span class="mi">30</span><span class="p">)</span>
    <span class="n">data</span> <span class="o">=</span> <span class="n">resp</span><span class="p">.</span><span class="n">json</span><span class="p">()</span>
    <span class="k">if</span> <span class="nb">isinstance</span><span class="p">(</span><span class="n">data</span><span class="p">,</span> <span class="nb">dict</span><span class="p">)</span> <span class="ow">and</span> <span class="n">data</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"ok"</span><span class="p">)</span> <span class="ow">is</span> <span class="bp">False</span><span class="p">:</span>
        <span class="k">return</span> <span class="bp">None</span>  <span class="c1"># no such site on this host
</span>    <span class="k">return</span> <span class="n">data</span>

<span class="n">jobs</span> <span class="o">=</span> <span class="n">lever_postings</span><span class="p">(</span><span class="s">"palantir"</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="nb">len</span><span class="p">(</span><span class="n">jobs</span><span class="p">),</span> <span class="s">"open postings"</span><span class="p">)</span>
<span class="k">for</span> <span class="n">p</span> <span class="ow">in</span> <span class="n">jobs</span><span class="p">[:</span><span class="mi">5</span><span class="p">]:</span>
    <span class="n">created</span> <span class="o">=</span> <span class="n">datetime</span><span class="p">.</span><span class="n">fromtimestamp</span><span class="p">(</span><span class="n">p</span><span class="p">[</span><span class="s">"createdAt"</span><span class="p">]</span> <span class="o">/</span> <span class="mi">1000</span><span class="p">,</span> <span class="n">tz</span><span class="o">=</span><span class="n">timezone</span><span class="p">.</span><span class="n">utc</span><span class="p">).</span><span class="n">date</span><span class="p">()</span>
    <span class="n">cats</span> <span class="o">=</span> <span class="n">p</span><span class="p">[</span><span class="s">"categories"</span><span class="p">]</span>
    <span class="k">print</span><span class="p">(</span><span class="n">p</span><span class="p">[</span><span class="s">"text"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">cats</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"team"</span><span class="p">),</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">cats</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"location"</span><span class="p">),</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">p</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"workplaceType"</span><span class="p">),</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">created</span><span class="p">)</span>
</code></pre></div></div>

<p>On the day I ran it, Palantir’s site returned 321 postings in one response.</p>

<h2 id="what-each-posting-contains">What each posting contains</h2>

<p>The fields you will use most:</p>

<table>
  <thead>
    <tr>
      <th>Field</th>
      <th>What it is</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">id</code></td>
      <td>Posting id, stable while the job is open</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">text</code></td>
      <td>The job title</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">categories</code></td>
      <td><code class="language-plaintext highlighter-rouge">team</code>, <code class="language-plaintext highlighter-rouge">department</code>, <code class="language-plaintext highlighter-rouge">location</code>, <code class="language-plaintext highlighter-rouge">commitment</code> (Full-time and so on) and <code class="language-plaintext highlighter-rouge">allLocations</code></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">workplaceType</code></td>
      <td><code class="language-plaintext highlighter-rouge">on-site</code>, <code class="language-plaintext highlighter-rouge">hybrid</code>, <code class="language-plaintext highlighter-rouge">remote</code> or <code class="language-plaintext highlighter-rouge">unspecified</code></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">country</code></td>
      <td>Two-letter country code, when set</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">createdAt</code></td>
      <td>Creation time in milliseconds since 1970</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">hostedUrl</code>, <code class="language-plaintext highlighter-rouge">applyUrl</code></td>
      <td>The public job page and its apply form</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">descriptionPlain</code>, <code class="language-plaintext highlighter-rouge">lists</code>, <code class="language-plaintext highlighter-rouge">additionalPlain</code></td>
      <td>The description in plain text and bullet lists</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">salaryRange</code></td>
      <td>Optional object with <code class="language-plaintext highlighter-rouge">currency</code>, <code class="language-plaintext highlighter-rouge">interval</code>, <code class="language-plaintext highlighter-rouge">min</code> and <code class="language-plaintext highlighter-rouge">max</code></td>
    </tr>
  </tbody>
</table>

<p><code class="language-plaintext highlighter-rouge">salaryRange</code> is optional, and in my checks many companies leave it empty, including Palantir and Spotify on that day. If pay matters to you, also search <code class="language-plaintext highlighter-rouge">descriptionPlain</code> and <code class="language-plaintext highlighter-rouge">additionalPlain</code>, where some companies write the range as text.</p>

<h2 id="paging-filters-and-politeness">Paging, filters and politeness</h2>

<p>The endpoint accepts <code class="language-plaintext highlighter-rouge">skip</code> and <code class="language-plaintext highlighter-rouge">limit</code>, plus filters such as <code class="language-plaintext highlighter-rouge">team</code>, <code class="language-plaintext highlighter-rouge">location</code>, <code class="language-plaintext highlighter-rouge">commitment</code> and <code class="language-plaintext highlighter-rouge">department</code> (case sensitive when you pass several values). Without <code class="language-plaintext highlighter-rouge">limit</code> it returns the full list, which is usually what you want. Lever’s docs describe a rate limit for application POSTs, not for reading postings, but the same manners apply: one request per site per run, cached, with a clear User-Agent.</p>

<h2 id="where-the-one-script-approach-runs-out">Where the one-script approach runs out</h2>

<p>The script above is fine for a few companies. At 50 or 500 companies you run into four jobs you did not plan for:</p>

<ol>
  <li><strong>Finding the site name.</strong> Your list is company names and websites, not Lever site slugs. Some companies are on the EU host, and many are not on Lever at all.</li>
  <li><strong>Mixed systems.</strong> Sales and recruiting lists mix Lever with Greenhouse, Ashby, Workday and smaller European systems.</li>
  <li><strong>Change tracking.</strong> Lever shows what is open. New and closed jobs since last week need stored state and a diff.</li>
  <li><strong>One schema.</strong> Location, seniority, function and salary need normalizing before a spreadsheet or CRM can use them.</li>
</ol>

<h2 id="a-hosted-option">A hosted option</h2>

<p>I built <a href="https://apify.com/conserving_celerytop/lever-jobs-api">Lever Jobs API</a> on Apify for that case. It reads the same public Postings API, handles both regions, and returns one row per job in a fixed format. You pass Lever links, websites or company names.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"apify-client&gt;=3.2,&lt;4"</span>
<span class="nb">export </span><span class="nv">APIFY_TOKEN</span><span class="o">=</span>&lt;YOUR_APIFY_TOKEN&gt;
</code></pre></div></div>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">os</span>
<span class="kn">from</span> <span class="nn">decimal</span> <span class="kn">import</span> <span class="n">Decimal</span>
<span class="kn">from</span> <span class="nn">apify_client</span> <span class="kn">import</span> <span class="n">ApifyClient</span>

<span class="n">client</span> <span class="o">=</span> <span class="n">ApifyClient</span><span class="p">(</span><span class="n">os</span><span class="p">.</span><span class="n">environ</span><span class="p">[</span><span class="s">"APIFY_TOKEN"</span><span class="p">])</span>

<span class="n">run</span> <span class="o">=</span> <span class="n">client</span><span class="p">.</span><span class="n">actor</span><span class="p">(</span><span class="s">"conserving_celerytop/lever-jobs-api"</span><span class="p">).</span><span class="n">call</span><span class="p">(</span>
    <span class="n">run_input</span><span class="o">=</span><span class="p">{</span>
        <span class="s">"companies"</span><span class="p">:</span> <span class="p">[</span><span class="s">"https://jobs.lever.co/palantir"</span><span class="p">,</span> <span class="s">"https://jobs.lever.co/spotify"</span><span class="p">],</span>
        <span class="s">"jobFunctions"</span><span class="p">:</span> <span class="p">[</span><span class="s">"engineering"</span><span class="p">,</span> <span class="s">"data"</span><span class="p">],</span>
        <span class="s">"postedSince"</span><span class="p">:</span> <span class="s">"14 days"</span><span class="p">,</span>
    <span class="p">},</span>
    <span class="n">max_total_charge_usd</span><span class="o">=</span><span class="n">Decimal</span><span class="p">(</span><span class="s">"0.50"</span><span class="p">),</span>
<span class="p">)</span>

<span class="n">rows</span> <span class="o">=</span> <span class="nb">list</span><span class="p">(</span><span class="n">client</span><span class="p">.</span><span class="n">dataset</span><span class="p">(</span><span class="n">run</span><span class="p">.</span><span class="n">default_dataset_id</span><span class="p">).</span><span class="n">iterate_items</span><span class="p">())</span>
<span class="n">jobs</span> <span class="o">=</span> <span class="p">[</span><span class="n">r</span> <span class="k">for</span> <span class="n">r</span> <span class="ow">in</span> <span class="n">rows</span> <span class="k">if</span> <span class="n">r</span><span class="p">[</span><span class="s">"rowType"</span><span class="p">]</span> <span class="o">==</span> <span class="s">"job"</span><span class="p">]</span>
<span class="k">print</span><span class="p">(</span><span class="nb">len</span><span class="p">(</span><span class="n">jobs</span><span class="p">),</span> <span class="s">"matching jobs"</span><span class="p">)</span>
<span class="k">for</span> <span class="n">r</span> <span class="ow">in</span> <span class="n">jobs</span><span class="p">[:</span><span class="mi">10</span><span class="p">]:</span>
    <span class="k">print</span><span class="p">(</span><span class="n">r</span><span class="p">[</span><span class="s">"companySlug"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">r</span><span class="p">[</span><span class="s">"title"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">r</span><span class="p">[</span><span class="s">"seniority"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">r</span><span class="p">[</span><span class="s">"location"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">r</span><span class="p">[</span><span class="s">"url"</span><span class="p">])</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">createdAt</code> becomes an ISO date in <code class="language-plaintext highlighter-rouge">postedAt</code>. <code class="language-plaintext highlighter-rouge">categories.location</code> becomes <code class="language-plaintext highlighter-rouge">location</code> plus <code class="language-plaintext highlighter-rouge">countryCode</code>. The title is read for <code class="language-plaintext highlighter-rouge">seniority</code> (intern to C-level) and <code class="language-plaintext highlighter-rouge">jobFunction</code> (engineering, sales, data and so on), so you can filter the rows before they reach your spreadsheet. A company that cannot be found gets a status row that says why, and that row is not dropped.</p>

<p>For a watchlist, add <code class="language-plaintext highlighter-rouge">"onlyNewJobs": true</code> and a <code class="language-plaintext highlighter-rouge">"monitorName"</code>. After the first run, each run returns only new and closed postings. Save it as an Apify task and schedule it daily or weekly.</p>

<h2 id="pricing">Pricing</h2>

<p>Free plan prices on the Apify Store on September 27, 2026. Check the Store page before you rely on them.</p>

<ul>
  <li><strong>Lever Jobs API:</strong> $0.10 per company with up to 1,000 open jobs ($0.09 on Scale, $0.07 on Business), then $0.045 per further 1,000. A later <code class="language-plaintext highlighter-rouge">onlyNewJobs</code> check is $0.002 per 1,000 open jobs. 100 Lever companies cost $10 for a full pull and about $0.20 per daily check after that. The example above, two companies, costs $0.20.</li>
  <li><strong><a href="https://apify.com/conserving_celerytop/live-career-page-jobs-api">ATS Jobs API</a></strong> reads Lever and 21 other systems with the same fields. It costs $0.045 per company with up to 1,000 jobs, less than the Lever-only version, so it is worth a look even for an all-Lever list.</li>
</ul>

<p>Paid Apify plans pay less per company. The free plan includes $5 of credit a month.</p>

<h2 id="when-you-should-not-use-it">When you should not use it</h2>

<ul>
  <li>A few Lever sites and no need for history: call the endpoint directly. It is free.</li>
  <li>You need every Lever job on the internet, searched by keyword: neither the endpoint nor the Actor does that. You supply the companies.</li>
  <li>You need candidates or pipeline data from your own Lever account: that is the authenticated Lever API, with a key from your admin.</li>
</ul>

<h2 id="faq">FAQ</h2>

<p><strong>Does the Lever Postings API need an API key?</strong>
Not for reading published postings. A key is needed to post applications and for the main Lever API.</p>

<p><strong>Why do I get “Document not found”?</strong>
The site name is wrong, the board is on the EU host, or the company does not use Lever.</p>

<p><strong>How do I get the posting date?</strong>
<code class="language-plaintext highlighter-rouge">createdAt</code> is milliseconds since 1970. Divide by 1,000 and convert to a date.</p>

<hr />

<p><em>Lever is a trademark of its owner. This article and the Actors are not affiliated with or endorsed by Lever. The Actors read only published postings, with no login.</em></p>]]></content><author><name>Don Mangu</name></author><category term="python" /><category term="api" /><category term="web-scraping" /><category term="apify" /><category term="recruitment" /><summary type="html"><![CDATA[Read job postings from any Lever board with the public Postings API in Python: endpoints, EU boards, dates, salary, and a hosted option for many boards.]]></summary></entry><entry><title type="html">Tech Stack Lookup API: CMS, Ecommerce and Analytics for Any Website</title><link href="https://donmangudata-ops.github.io/tech-stack-lookup-api-python/" rel="alternate" type="text/html" title="Tech Stack Lookup API: CMS, Ecommerce and Analytics for Any Website" /><published>2026-09-28T09:00:00+00:00</published><updated>2026-09-28T09:00:00+00:00</updated><id>https://donmangudata-ops.github.io/tech-stack-lookup-api-python</id><content type="html" xml:base="https://donmangudata-ops.github.io/tech-stack-lookup-api-python/"><![CDATA[<p><em>Disclosure: I built the Apify Actor in the second half of this article, and it is paid. This article was drafted with an AI assistant. Prices and field names were checked on September 27, 2026.</em></p>

<p>Knowing what a website runs on is useful in a lot of jobs. An agency that builds on Shopify wants the Shopify stores in its market. A payments company wants to know who already uses Stripe. A sales team wants to sort 2,000 leads by CMS before it writes a single email.</p>

<p>You can do this by hand with a browser extension, one site at a time. For a list, you want an API. This guide explains how detection works, gives you a small script to try it yourself, and then shows a hosted option for large lists.</p>

<h2 id="how-tech-stack-detection-works">How tech stack detection works</h2>

<p>Most tools, commercial or open source, work the same way. They load a page and look for fingerprints:</p>

<ul>
  <li><strong>Script and stylesheet URLs.</strong> <code class="language-plaintext highlighter-rouge">cdn.shopify.com</code> means Shopify, <code class="language-plaintext highlighter-rouge">js.stripe.com</code> means Stripe, <code class="language-plaintext highlighter-rouge">/wp-content/</code> means WordPress.</li>
  <li><strong>HTML markers.</strong> A <code class="language-plaintext highlighter-rouge">&lt;meta name="generator"&gt;</code> tag, or attributes such as <code class="language-plaintext highlighter-rouge">data-wf-site</code> on Webflow sites.</li>
  <li><strong>Response headers and cookies.</strong> <code class="language-plaintext highlighter-rouge">Server</code>, <code class="language-plaintext highlighter-rouge">X-Powered-By</code> and cookie names often name the platform or host.</li>
  <li><strong>Versions</strong>, when a file name or tag includes one.</li>
</ul>

<p>A widely used open fingerprint set started in the Wappalyzer project and is now maintained as <a href="https://github.com/enthec/webappanalyzer">webappanalyzer</a>, with thousands of technologies and their patterns. Accuracy depends on what the page shows. A tool that is loaded later by JavaScript, for example an analytics tag inside a tag manager, only appears if you run the page in a browser.</p>

<h2 id="a-small-script-to-try-it-yourself">A small script to try it yourself</h2>

<p>This script checks a site’s <code class="language-plaintext highlighter-rouge">robots.txt</code>, loads the homepage once, and looks for a few fingerprints. The patterns are examples, not a full set.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">re</span>
<span class="kn">import</span> <span class="nn">requests</span>
<span class="kn">from</span> <span class="nn">urllib.robotparser</span> <span class="kn">import</span> <span class="n">RobotFileParser</span>

<span class="n">UA</span> <span class="o">=</span> <span class="s">"stack-check-example/0.1 (+https://example.com/contact)"</span>
<span class="n">SIGNS</span> <span class="o">=</span> <span class="p">{</span>
    <span class="s">"Shopify"</span><span class="p">:</span> <span class="p">[</span><span class="sa">r</span><span class="s">"cdn\.shopify\.com"</span><span class="p">,</span> <span class="sa">r</span><span class="s">"Shopify\.theme"</span><span class="p">],</span>
    <span class="s">"WordPress"</span><span class="p">:</span> <span class="p">[</span><span class="sa">r</span><span class="s">"/wp-content/"</span><span class="p">,</span> <span class="sa">r</span><span class="s">"/wp-includes/"</span><span class="p">],</span>
    <span class="s">"Webflow"</span><span class="p">:</span> <span class="p">[</span><span class="sa">r</span><span class="s">"data-wf-site"</span><span class="p">,</span> <span class="sa">r</span><span class="s">"webflow\.js"</span><span class="p">],</span>
    <span class="s">"Google Analytics"</span><span class="p">:</span> <span class="p">[</span><span class="sa">r</span><span class="s">"googletagmanager\.com/gtag/js"</span><span class="p">,</span> <span class="sa">r</span><span class="s">"google-analytics\.com/analytics\.js"</span><span class="p">],</span>
    <span class="s">"Google Tag Manager"</span><span class="p">:</span> <span class="p">[</span><span class="sa">r</span><span class="s">"googletagmanager\.com/gtm\.js"</span><span class="p">],</span>
    <span class="s">"HubSpot"</span><span class="p">:</span> <span class="p">[</span><span class="sa">r</span><span class="s">"js\.hs-scripts\.com"</span><span class="p">,</span> <span class="sa">r</span><span class="s">"js\.hsforms\.net"</span><span class="p">],</span>
    <span class="s">"Stripe"</span><span class="p">:</span> <span class="p">[</span><span class="sa">r</span><span class="s">"js\.stripe\.com"</span><span class="p">],</span>
<span class="p">}</span>

<span class="k">def</span> <span class="nf">check</span><span class="p">(</span><span class="n">domain</span><span class="p">):</span>
    <span class="n">base</span> <span class="o">=</span> <span class="sa">f</span><span class="s">"https://</span><span class="si">{</span><span class="n">domain</span><span class="si">}</span><span class="s">"</span>
    <span class="n">robots</span> <span class="o">=</span> <span class="n">RobotFileParser</span><span class="p">()</span>
    <span class="n">r</span> <span class="o">=</span> <span class="n">requests</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="n">base</span> <span class="o">+</span> <span class="s">"/robots.txt"</span><span class="p">,</span> <span class="n">headers</span><span class="o">=</span><span class="p">{</span><span class="s">"User-Agent"</span><span class="p">:</span> <span class="n">UA</span><span class="p">},</span> <span class="n">timeout</span><span class="o">=</span><span class="mi">20</span><span class="p">)</span>
    <span class="n">robots</span><span class="p">.</span><span class="n">parse</span><span class="p">(</span><span class="n">r</span><span class="p">.</span><span class="n">text</span><span class="p">.</span><span class="n">splitlines</span><span class="p">()</span> <span class="k">if</span> <span class="n">r</span><span class="p">.</span><span class="n">status_code</span> <span class="o">==</span> <span class="mi">200</span> <span class="k">else</span> <span class="p">[])</span>
    <span class="k">if</span> <span class="ow">not</span> <span class="n">robots</span><span class="p">.</span><span class="n">can_fetch</span><span class="p">(</span><span class="n">UA</span><span class="p">,</span> <span class="n">base</span> <span class="o">+</span> <span class="s">"/"</span><span class="p">):</span>
        <span class="k">return</span> <span class="p">{</span><span class="s">"domain"</span><span class="p">:</span> <span class="n">domain</span><span class="p">,</span> <span class="s">"status"</span><span class="p">:</span> <span class="s">"robots_disallowed"</span><span class="p">,</span> <span class="s">"found"</span><span class="p">:</span> <span class="p">[]}</span>
    <span class="n">resp</span> <span class="o">=</span> <span class="n">requests</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="n">base</span> <span class="o">+</span> <span class="s">"/"</span><span class="p">,</span> <span class="n">headers</span><span class="o">=</span><span class="p">{</span><span class="s">"User-Agent"</span><span class="p">:</span> <span class="n">UA</span><span class="p">},</span> <span class="n">timeout</span><span class="o">=</span><span class="mi">20</span><span class="p">)</span>
    <span class="n">html</span> <span class="o">=</span> <span class="n">resp</span><span class="p">.</span><span class="n">text</span>
    <span class="n">found</span> <span class="o">=</span> <span class="p">[</span><span class="n">name</span> <span class="k">for</span> <span class="n">name</span><span class="p">,</span> <span class="n">pats</span> <span class="ow">in</span> <span class="n">SIGNS</span><span class="p">.</span><span class="n">items</span><span class="p">()</span> <span class="k">if</span> <span class="nb">any</span><span class="p">(</span><span class="n">re</span><span class="p">.</span><span class="n">search</span><span class="p">(</span><span class="n">p</span><span class="p">,</span> <span class="n">html</span><span class="p">)</span> <span class="k">for</span> <span class="n">p</span> <span class="ow">in</span> <span class="n">pats</span><span class="p">)]</span>
    <span class="n">gen</span> <span class="o">=</span> <span class="n">re</span><span class="p">.</span><span class="n">search</span><span class="p">(</span><span class="sa">r</span><span class="s">'&lt;meta[^&gt;]+name=["\']generator["\'][^&gt;]+content=["\']([^"\']+)'</span><span class="p">,</span> <span class="n">html</span><span class="p">,</span> <span class="n">re</span><span class="p">.</span><span class="n">I</span><span class="p">)</span>
    <span class="k">if</span> <span class="n">gen</span><span class="p">:</span>
        <span class="n">found</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="s">"generator: "</span> <span class="o">+</span> <span class="n">gen</span><span class="p">.</span><span class="n">group</span><span class="p">(</span><span class="mi">1</span><span class="p">))</span>
    <span class="k">return</span> <span class="p">{</span><span class="s">"domain"</span><span class="p">:</span> <span class="n">domain</span><span class="p">,</span> <span class="s">"status"</span><span class="p">:</span> <span class="n">resp</span><span class="p">.</span><span class="n">status_code</span><span class="p">,</span> <span class="s">"server"</span><span class="p">:</span> <span class="n">resp</span><span class="p">.</span><span class="n">headers</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="s">"server"</span><span class="p">),</span> <span class="s">"found"</span><span class="p">:</span> <span class="n">found</span><span class="p">}</span>

<span class="k">print</span><span class="p">(</span><span class="n">check</span><span class="p">(</span><span class="s">"pypi.org"</span><span class="p">))</span>
</code></pre></div></div>

<p>Put your own contact link in the User-Agent. Site owners appreciate knowing who is calling.</p>

<p>This works for a handful of sites. At a few thousand, the fingerprint list becomes the real work: keeping patterns current, scoring confidence, reading versions, and handling redirects, timeouts and blocked sites without losing track of which is which.</p>

<h2 id="a-hosted-option-website-tech-stack-detector">A hosted option: Website Tech Stack Detector</h2>

<p><a href="https://apify.com/conserving_celerytop/website-tech-stack-detector">Website Tech Stack Detector</a> is an Apify Actor I built for lists. It uses the open fingerprint set, more than 7,600 technologies, and makes one homepage request per site, only when that site’s <code class="language-plaintext highlighter-rouge">robots.txt</code> allows it. You send up to 10,000 domains or URLs per run.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"apify-client&gt;=3.2,&lt;4"</span>
<span class="nb">export </span><span class="nv">APIFY_TOKEN</span><span class="o">=</span>&lt;YOUR_APIFY_TOKEN&gt;
</code></pre></div></div>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">csv</span><span class="p">,</span> <span class="n">os</span>
<span class="kn">from</span> <span class="nn">decimal</span> <span class="kn">import</span> <span class="n">Decimal</span>
<span class="kn">from</span> <span class="nn">apify_client</span> <span class="kn">import</span> <span class="n">ApifyClient</span>

<span class="n">client</span> <span class="o">=</span> <span class="n">ApifyClient</span><span class="p">(</span><span class="n">os</span><span class="p">.</span><span class="n">environ</span><span class="p">[</span><span class="s">"APIFY_TOKEN"</span><span class="p">])</span>

<span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="s">"sites.txt"</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
    <span class="n">sites</span> <span class="o">=</span> <span class="p">[</span><span class="n">line</span><span class="p">.</span><span class="n">strip</span><span class="p">()</span> <span class="k">for</span> <span class="n">line</span> <span class="ow">in</span> <span class="n">f</span> <span class="k">if</span> <span class="n">line</span><span class="p">.</span><span class="n">strip</span><span class="p">()]</span>

<span class="n">run</span> <span class="o">=</span> <span class="n">client</span><span class="p">.</span><span class="n">actor</span><span class="p">(</span><span class="s">"conserving_celerytop/website-tech-stack-detector"</span><span class="p">).</span><span class="n">call</span><span class="p">(</span>
    <span class="n">run_input</span><span class="o">=</span><span class="p">{</span><span class="s">"websites"</span><span class="p">:</span> <span class="n">sites</span><span class="p">,</span> <span class="s">"maxSites"</span><span class="p">:</span> <span class="nb">len</span><span class="p">(</span><span class="n">sites</span><span class="p">),</span> <span class="s">"categories"</span><span class="p">:</span> <span class="p">[</span><span class="s">"CMS"</span><span class="p">,</span> <span class="s">"Ecommerce"</span><span class="p">,</span> <span class="s">"Analytics"</span><span class="p">]},</span>
    <span class="n">max_total_charge_usd</span><span class="o">=</span><span class="n">Decimal</span><span class="p">(</span><span class="s">"5.00"</span><span class="p">),</span>
<span class="p">)</span>

<span class="n">cols</span> <span class="o">=</span> <span class="p">[</span><span class="s">"domain"</span><span class="p">,</span> <span class="s">"status"</span><span class="p">,</span> <span class="s">"cms"</span><span class="p">,</span> <span class="s">"ecommerce"</span><span class="p">,</span> <span class="s">"analytics"</span><span class="p">,</span> <span class="s">"paymentProcessors"</span><span class="p">]</span>
<span class="k">with</span> <span class="nb">open</span><span class="p">(</span><span class="s">"tech_stack.csv"</span><span class="p">,</span> <span class="s">"w"</span><span class="p">,</span> <span class="n">newline</span><span class="o">=</span><span class="s">""</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
    <span class="n">w</span> <span class="o">=</span> <span class="n">csv</span><span class="p">.</span><span class="n">writer</span><span class="p">(</span><span class="n">f</span><span class="p">)</span>
    <span class="n">w</span><span class="p">.</span><span class="n">writerow</span><span class="p">(</span><span class="n">cols</span><span class="p">)</span>
    <span class="k">for</span> <span class="n">row</span> <span class="ow">in</span> <span class="n">client</span><span class="p">.</span><span class="n">dataset</span><span class="p">(</span><span class="n">run</span><span class="p">.</span><span class="n">default_dataset_id</span><span class="p">).</span><span class="n">iterate_items</span><span class="p">():</span>
        <span class="n">w</span><span class="p">.</span><span class="n">writerow</span><span class="p">([</span><span class="n">row</span><span class="p">[</span><span class="s">"domain"</span><span class="p">],</span> <span class="n">row</span><span class="p">[</span><span class="s">"status"</span><span class="p">]]</span> <span class="o">+</span> <span class="p">[</span><span class="s">", "</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="n">row</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="n">c</span><span class="p">)</span> <span class="ow">or</span> <span class="p">[])</span> <span class="k">for</span> <span class="n">c</span> <span class="ow">in</span> <span class="n">cols</span><span class="p">[</span><span class="mi">2</span><span class="p">:]])</span>
</code></pre></div></div>

<p>Leave out <code class="language-plaintext highlighter-rouge">categories</code> to get everything it finds. Each row has:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">technologies</code>: every match with <code class="language-plaintext highlighter-rouge">name</code>, <code class="language-plaintext highlighter-rouge">version</code> when the page shows one, <code class="language-plaintext highlighter-rouge">confidence</code> from 0 to 100 and <code class="language-plaintext highlighter-rouge">categories</code></li>
  <li>quick columns ready for a spreadsheet: <code class="language-plaintext highlighter-rouge">cms</code>, <code class="language-plaintext highlighter-rouge">ecommerce</code>, <code class="language-plaintext highlighter-rouge">analytics</code>, <code class="language-plaintext highlighter-rouge">frameworks</code>, <code class="language-plaintext highlighter-rouge">cdn</code>, <code class="language-plaintext highlighter-rouge">hosting</code>, <code class="language-plaintext highlighter-rouge">paymentProcessors</code>, <code class="language-plaintext highlighter-rouge">tagManagers</code></li>
  <li><code class="language-plaintext highlighter-rouge">status</code>: <code class="language-plaintext highlighter-rouge">ok</code>, or the reason a site was not checked, such as <code class="language-plaintext highlighter-rouge">robots_disallowed</code>, <code class="language-plaintext highlighter-rouge">blocked</code>, <code class="language-plaintext highlighter-rouge">timeout</code> or <code class="language-plaintext highlighter-rouge">domain_not_found</code></li>
</ul>

<p>The same run also works over plain HTTP with Apify’s <code class="language-plaintext highlighter-rouge">run-sync-get-dataset-items</code> endpoint, if you would rather not install the client.</p>

<h2 id="what-it-costs">What it costs</h2>

<p>Store prices on September 27, 2026. Check the Store page before relying on them.</p>

<ul>
  <li>$0.002 per website whose homepage loaded, which is $2 per 1,000, on the free and Starter plans. Scale pays $0.0018 and Business $0.0016.</li>
  <li>Apify’s start event adds $0.00005 per run for each GB of memory.</li>
  <li>Sites closed by <code class="language-plaintext highlighter-rouge">robots.txt</code>, blocked, timed out or not found cost nothing, and neither do duplicates or invalid entries.</li>
  <li>A site that loads but shows no known technology is charged, since the page was fetched and checked.</li>
  <li>Example: 500 domains where 460 load cost $0.92.</li>
</ul>

<h2 id="honest-limits">Honest limits</h2>

<ul>
  <li><strong>One page, no JavaScript.</strong> It reads the homepage HTML and headers the server sends. Tools injected later by scripts can be missed. If you need those, run a headless browser, which costs more per site.</li>
  <li><strong>Homepage only.</strong> A checkout platform used only on <code class="language-plaintext highlighter-rouge">shop.example.com</code> will not show on <code class="language-plaintext highlighter-rouge">example.com</code>.</li>
  <li><strong>No usage history.</strong> It tells you what a site runs today, not when it switched or what it spends. Commercial databases that crawl the web continuously are the tool for “every site that used X since 2021”.</li>
  <li><strong>robots.txt is respected.</strong> Some large sites disallow all bots, and those come back as <code class="language-plaintext highlighter-rouge">robots_disallowed</code>, free.</li>
</ul>

<h2 id="faq">FAQ</h2>

<p><strong>Is there a free tech stack lookup API?</strong>
For a few sites, a browser extension or a script like the one above is free. Most hosted APIs charge per lookup or per month.</p>

<p><strong>Can I look up which companies use a technology?</strong>
Not directly. Send a list of candidate domains and filter the results by technology. The Actor does not keep a crawl of the whole web.</p>

<p><strong>Does it collect personal data?</strong>
No. It reads public homepages once per site and does not store cookie values or personal data.</p>

<hr />

<p><em>This article and the Actor are not affiliated with or endorsed by any company whose technology it detects. The fingerprint data is open source under GPL-3.0.</em></p>]]></content><author><name>Don Mangu</name></author><category term="python" /><category term="api" /><category term="web-scraping" /><category term="webdev" /><category term="apify" /><summary type="html"><![CDATA[Look up the CMS, ecommerce, analytics and payment tools of many websites in Python: how detection works, a small DIY script, and a $2 per 1,000 API.]]></summary></entry><entry><title type="html">Greenhouse Jobs API: Get Every Open Job From a Board in Python</title><link href="https://donmangudata-ops.github.io/greenhouse-jobs-api-python/" rel="alternate" type="text/html" title="Greenhouse Jobs API: Get Every Open Job From a Board in Python" /><published>2026-09-27T09:00:00+00:00</published><updated>2026-09-27T09:00:00+00:00</updated><id>https://donmangudata-ops.github.io/greenhouse-jobs-api-python</id><content type="html" xml:base="https://donmangudata-ops.github.io/greenhouse-jobs-api-python/"><![CDATA[<p><em>Disclosure: I built the Apify Actors mentioned near the end, and they are paid. This article was drafted with an AI assistant. The endpoints were checked against live boards on September 26, 2026. Actor prices and field names were checked on September 27, 2026.</em></p>

<p>Greenhouse is the applicant tracking system behind the careers pages of a lot of tech companies. If you want those jobs as data, for a job site, a recruiting list or a sales signal, you do not need a Greenhouse account. Every public Greenhouse board has a read-only Job Board API that answers without a key.</p>

<p>This guide covers that free endpoint first, then what gets hard once you have more than a handful of companies.</p>

<h2 id="step-1-find-the-board-token">Step 1: find the board token</h2>

<p>Each Greenhouse board has a short name, called the board token. You can see it in the board link:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">https://boards.greenhouse.io/stripe</code> has the token <code class="language-plaintext highlighter-rouge">stripe</code></li>
  <li><code class="language-plaintext highlighter-rouge">https://job-boards.greenhouse.io/stripe</code> is the newer host, same token</li>
</ul>

<p>Many companies show their jobs on their own site instead. Open any job there and look at the link. A <code class="language-plaintext highlighter-rouge">gh_jid=</code> parameter means the page embeds a Greenhouse board, and the page source usually loads a <code class="language-plaintext highlighter-rouge">boards.greenhouse.io/embed/job_board</code> script whose <code class="language-plaintext highlighter-rouge">for=</code> parameter is the token.</p>

<h2 id="step-2-call-the-job-board-api">Step 2: call the Job Board API</h2>

<p>The list endpoint is <code class="language-plaintext highlighter-rouge">https://boards-api.greenhouse.io/v1/boards/{token}/jobs</code>. It returns every published job in one response, with no paging.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">requests</span>

<span class="n">token</span> <span class="o">=</span> <span class="s">"stripe"</span>
<span class="n">url</span> <span class="o">=</span> <span class="sa">f</span><span class="s">"https://boards-api.greenhouse.io/v1/boards/</span><span class="si">{</span><span class="n">token</span><span class="si">}</span><span class="s">/jobs"</span>
<span class="n">data</span> <span class="o">=</span> <span class="n">requests</span><span class="p">.</span><span class="n">get</span><span class="p">(</span><span class="n">url</span><span class="p">,</span> <span class="n">params</span><span class="o">=</span><span class="p">{</span><span class="s">"content"</span><span class="p">:</span> <span class="s">"true"</span><span class="p">},</span> <span class="n">timeout</span><span class="o">=</span><span class="mi">30</span><span class="p">).</span><span class="n">json</span><span class="p">()</span>

<span class="k">print</span><span class="p">(</span><span class="n">data</span><span class="p">[</span><span class="s">"meta"</span><span class="p">][</span><span class="s">"total"</span><span class="p">],</span> <span class="s">"open jobs"</span><span class="p">)</span>
<span class="k">for</span> <span class="n">job</span> <span class="ow">in</span> <span class="n">data</span><span class="p">[</span><span class="s">"jobs"</span><span class="p">][:</span><span class="mi">5</span><span class="p">]:</span>
    <span class="k">print</span><span class="p">(</span><span class="n">job</span><span class="p">[</span><span class="s">"title"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">job</span><span class="p">[</span><span class="s">"location"</span><span class="p">][</span><span class="s">"name"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">job</span><span class="p">[</span><span class="s">"first_published"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">job</span><span class="p">[</span><span class="s">"absolute_url"</span><span class="p">])</span>
</code></pre></div></div>

<p>When I ran this, Stripe’s board returned 702 jobs. Each job has <code class="language-plaintext highlighter-rouge">id</code>, <code class="language-plaintext highlighter-rouge">title</code>, <code class="language-plaintext highlighter-rouge">location.name</code>, <code class="language-plaintext highlighter-rouge">first_published</code>, <code class="language-plaintext highlighter-rouge">updated_at</code>, <code class="language-plaintext highlighter-rouge">requisition_id</code> and <code class="language-plaintext highlighter-rouge">absolute_url</code>. With <code class="language-plaintext highlighter-rouge">content=true</code> you also get <code class="language-plaintext highlighter-rouge">content</code> (the description as escaped HTML), <code class="language-plaintext highlighter-rouge">departments</code> and <code class="language-plaintext highlighter-rouge">offices</code>.</p>

<p>Two details that trip people up:</p>

<ol>
  <li><code class="language-plaintext highlighter-rouge">content</code> is HTML with entities escaped, so run it through <code class="language-plaintext highlighter-rouge">html.unescape</code> before you parse it.</li>
  <li>Pay ranges are not in the list. Request a single job with <code class="language-plaintext highlighter-rouge">/jobs/{id}?pay_transparency=true</code> and read <code class="language-plaintext highlighter-rouge">pay_input_ranges</code>. It is often an empty list, because many companies publish pay only in the description text.</li>
</ol>

<p>That is all you need for one company. The API is public, documented by Greenhouse and fast. If you track two or three boards, stop here and write a small cron job.</p>

<h2 id="where-it-gets-harder">Where it gets harder</h2>

<p>The work grows once your list grows.</p>

<ul>
  <li><strong>Finding tokens.</strong> A list of 300 company websites does not come with board tokens. Some redirect to Greenhouse, some embed it, and some use a different system entirely.</li>
  <li><strong>Other systems.</strong> In most real account lists, a good share of companies run Lever, Ashby, Workday or a European ATS such as Personio or Teamtailor. Each has its own endpoint and field names.</li>
  <li><strong>New and closed jobs.</strong> The API shows what is open now. To see what changed since yesterday you have to store every run and diff it yourself.</li>
  <li><strong>Normalizing.</strong> Location strings, remote flags, seniority and salary all come in different shapes.</li>
</ul>

<p>You can build all of that. I did, and it took far longer than the first script.</p>

<h2 id="a-hosted-option-for-many-companies">A hosted option for many companies</h2>

<p>I packaged that work as an Apify Actor called <a href="https://apify.com/conserving_celerytop/greenhouse-jobs-api">Greenhouse Jobs API</a>. You give it board links, company websites or plain names and it reads the same public Greenhouse endpoint for each one, then returns one row per job in a fixed format.</p>

<p>You need an Apify account and its API token (Console, Settings, API &amp; Integrations).</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"apify-client&gt;=3.2,&lt;4"</span>
<span class="nb">export </span><span class="nv">APIFY_TOKEN</span><span class="o">=</span>&lt;YOUR_APIFY_TOKEN&gt;
</code></pre></div></div>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">os</span>
<span class="kn">from</span> <span class="nn">decimal</span> <span class="kn">import</span> <span class="n">Decimal</span>
<span class="kn">from</span> <span class="nn">apify_client</span> <span class="kn">import</span> <span class="n">ApifyClient</span>

<span class="n">client</span> <span class="o">=</span> <span class="n">ApifyClient</span><span class="p">(</span><span class="n">os</span><span class="p">.</span><span class="n">environ</span><span class="p">[</span><span class="s">"APIFY_TOKEN"</span><span class="p">])</span>

<span class="n">run</span> <span class="o">=</span> <span class="n">client</span><span class="p">.</span><span class="n">actor</span><span class="p">(</span><span class="s">"conserving_celerytop/greenhouse-jobs-api"</span><span class="p">).</span><span class="n">call</span><span class="p">(</span>
    <span class="n">run_input</span><span class="o">=</span><span class="p">{</span>
        <span class="s">"companies"</span><span class="p">:</span> <span class="p">[</span>
            <span class="s">"https://boards.greenhouse.io/stripe"</span><span class="p">,</span>
            <span class="s">"https://job-boards.greenhouse.io/dropbox"</span><span class="p">,</span>
            <span class="s">"figma.com"</span><span class="p">,</span>
        <span class="p">],</span>
        <span class="s">"titleIncludes"</span><span class="p">:</span> <span class="p">[</span><span class="s">"engineer"</span><span class="p">],</span>
        <span class="s">"postedSince"</span><span class="p">:</span> <span class="s">"30 days"</span><span class="p">,</span>
    <span class="p">},</span>
    <span class="n">max_total_charge_usd</span><span class="o">=</span><span class="n">Decimal</span><span class="p">(</span><span class="s">"0.50"</span><span class="p">),</span>
<span class="p">)</span>

<span class="k">for</span> <span class="n">row</span> <span class="ow">in</span> <span class="n">client</span><span class="p">.</span><span class="n">dataset</span><span class="p">(</span><span class="n">run</span><span class="p">.</span><span class="n">default_dataset_id</span><span class="p">).</span><span class="n">iterate_items</span><span class="p">():</span>
    <span class="k">if</span> <span class="n">row</span><span class="p">[</span><span class="s">"rowType"</span><span class="p">]</span> <span class="o">==</span> <span class="s">"job"</span><span class="p">:</span>
        <span class="k">print</span><span class="p">(</span><span class="n">row</span><span class="p">[</span><span class="s">"companySlug"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">row</span><span class="p">[</span><span class="s">"title"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">row</span><span class="p">[</span><span class="s">"countryCode"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">row</span><span class="p">[</span><span class="s">"postedAt"</span><span class="p">])</span>
    <span class="k">elif</span> <span class="n">row</span><span class="p">[</span><span class="s">"rowType"</span><span class="p">]</span> <span class="o">==</span> <span class="s">"status"</span><span class="p">:</span>
        <span class="k">print</span><span class="p">(</span><span class="n">row</span><span class="p">[</span><span class="s">"company"</span><span class="p">],</span> <span class="s">"-&gt;"</span><span class="p">,</span> <span class="n">row</span><span class="p">[</span><span class="s">"companyStatus"</span><span class="p">])</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">max_total_charge_usd</code> is a hard cap on what the run can cost, which is worth setting while you test.</p>

<p>Every job row has the same fields: <code class="language-plaintext highlighter-rouge">title</code>, <code class="language-plaintext highlighter-rouge">department</code>, <code class="language-plaintext highlighter-rouge">location</code>, <code class="language-plaintext highlighter-rouge">countryCode</code>, <code class="language-plaintext highlighter-rouge">workplaceType</code>, <code class="language-plaintext highlighter-rouge">seniority</code>, <code class="language-plaintext highlighter-rouge">jobFunction</code>, <code class="language-plaintext highlighter-rouge">salaryMin</code>, <code class="language-plaintext highlighter-rouge">salaryMax</code>, <code class="language-plaintext highlighter-rouge">salaryCurrency</code>, <code class="language-plaintext highlighter-rouge">postedAt</code>, <code class="language-plaintext highlighter-rouge">url</code> and <code class="language-plaintext highlighter-rouge">jobKey</code>. <code class="language-plaintext highlighter-rouge">jobKey</code> stays the same between runs, so you can join runs on it. A company with nothing to return gets one status row, such as <code class="language-plaintext highlighter-rouge">not_found</code> or <code class="language-plaintext highlighter-rouge">no_open_jobs</code>, instead of silently disappearing.</p>

<h3 id="only-new-and-closed-jobs">Only new and closed jobs</h3>

<p>Set <code class="language-plaintext highlighter-rouge">"onlyNewJobs": true</code> and give the watchlist a <code class="language-plaintext highlighter-rouge">"monitorName"</code>. The first run returns everything. Each later run returns only jobs that are new since the last run (<code class="language-plaintext highlighter-rouge">change: "new"</code>) and jobs that closed (<code class="language-plaintext highlighter-rouge">change: "closed"</code>). Save the input as a task in Apify, add a daily schedule, and you have a new-jobs feed without storing anything yourself.</p>

<h2 id="what-it-costs">What it costs</h2>

<p>These are the free plan prices on the Apify Store on September 27, 2026. Check the Store page before you rely on them.</p>

<ul>
  <li><strong>Greenhouse Jobs API:</strong> $0.10 per company, which includes up to 1,000 of its open jobs ($0.09 on the Scale plan, $0.07 on Business). Each further 1,000 jobs of the same company is $0.045. A later check with <code class="language-plaintext highlighter-rouge">onlyNewJobs</code> is $0.002 per 1,000 open jobs on the board. So 100 companies cost $10 for the first full pull, and a daily <code class="language-plaintext highlighter-rouge">onlyNewJobs</code> check of them about $0.20. Descriptions cost nothing extra.</li>
  <li><strong><a href="https://apify.com/conserving_celerytop/live-career-page-jobs-api">ATS Jobs API</a></strong>, the multi-board version from the same author, reads Greenhouse plus Lever, Ashby, Workday and 18 other systems with the same input and output. It costs $0.045 per company with up to 1,000 jobs, less than the single-board version, so it is worth a look even for a Greenhouse-only list, and it is the one to use when your list mixes systems.</li>
</ul>

<p>There are single-system versions for <a href="https://apify.com/conserving_celerytop/lever-jobs-api">Lever</a>, <a href="https://apify.com/conserving_celerytop/ashby-jobs-api">Ashby</a> and <a href="https://apify.com/conserving_celerytop/workday-jobs-api">Workday</a> too. Paid Apify plans get lower per-company prices. Apify’s free plan includes $5 of monthly credit.</p>

<h2 id="when-not-to-use-it">When not to use it</h2>

<ul>
  <li><strong>One or two boards.</strong> The free endpoint above is simpler and costs nothing.</li>
  <li><strong>Search across all companies.</strong> Neither the endpoint nor the Actor searches every Greenhouse board by keyword. You bring the list of companies. For keyword search across a fixed set of 824 startups and tech companies there is a separate Actor, <a href="https://apify.com/conserving_celerytop/tech-jobs-search">Tech Jobs Search</a>.</li>
  <li><strong>Your own candidates or applications.</strong> That is Greenhouse’s Harvest API, which needs a key from the company’s Greenhouse account. The Job Board API only shows published jobs.</li>
</ul>

<h2 id="faq">FAQ</h2>

<p><strong>Do I need a Greenhouse API key to read job postings?</strong>
No. The Job Board API is read-only and public for published jobs. Keys are only for posting applications and for the Harvest API.</p>

<p><strong>Is there a rate limit?</strong>
Greenhouse does not publish one for the Job Board API. Be polite: one request per board per run, and cache the result.</p>

<p><strong>Can I get jobs posted months ago that are still open?</strong>
Yes. The API returns every open job, whatever its age. <code class="language-plaintext highlighter-rouge">first_published</code> tells you how old it is.</p>

<hr />

<p><em>Greenhouse is a trademark of its owner. This article and the Actors are not affiliated with or endorsed by Greenhouse. The Actors read only jobs that companies publish on public boards, with no login.</em></p>]]></content><author><name>Don Mangu</name></author><category term="python" /><category term="api" /><category term="web-scraping" /><category term="apify" /><category term="recruitment" /><summary type="html"><![CDATA[Pull every open job from a Greenhouse board in Python: the free public endpoint, what it returns, its limits, and a hosted option for many boards.]]></summary></entry><entry><title type="html">Workday Jobs API: Pull Open Jobs From Workday Career Sites</title><link href="https://donmangudata-ops.github.io/workday-jobs-api-python/" rel="alternate" type="text/html" title="Workday Jobs API: Pull Open Jobs From Workday Career Sites" /><published>2026-09-27T09:00:00+00:00</published><updated>2026-09-27T09:00:00+00:00</updated><id>https://donmangudata-ops.github.io/workday-jobs-api-python</id><content type="html" xml:base="https://donmangudata-ops.github.io/workday-jobs-api-python/"><![CDATA[<p><em>Disclosure: I built the Apify Actor used in this guide, and it is paid. This article was drafted with an AI assistant. Field names and prices were checked on September 27, 2026.</em></p>

<p>Many large employers run their careers pages on Workday. Intel, Adobe and Salesforce are among them. If you searched for “Workday jobs API”, you probably want those open jobs as data, for a job board, a recruiting list or a sales signal. This guide covers what Workday offers outsiders, why a do-it-yourself scraper is harder than it looks, and a hosted option priced per company.</p>

<h2 id="two-different-things-called-workday-api">Two different things called “Workday API”</h2>

<p><strong>Workday’s official APIs</strong> (REST and SOAP Web Services) belong to the employer. You need an integration account inside that company’s Workday tenant, set up by its HR or IT team. They fit if you work there and want your own requisitions, not for reading other companies’ jobs.</p>

<p><strong>Public career sites</strong> are what job seekers see, at addresses like <code class="language-plaintext highlighter-rouge">https://intel.wd1.myworkdayjobs.com/External</code>. Anyone can open them without logging in, but Workday publishes no documented jobs API for outsiders.</p>

<h2 id="how-a-workday-career-site-link-is-built">How a Workday career site link is built</h2>

<p>Every link has three parts you need:</p>

<ul>
  <li>the <strong>tenant</strong>, the company’s name in Workday (<code class="language-plaintext highlighter-rouge">intel</code>)</li>
  <li>the <strong>pod</strong>, a data center label such as <code class="language-plaintext highlighter-rouge">wd1</code>, <code class="language-plaintext highlighter-rouge">wd5</code> or <code class="language-plaintext highlighter-rouge">wd12</code></li>
  <li>the <strong>site name</strong>, the path after the host (<code class="language-plaintext highlighter-rouge">External</code>, <code class="language-plaintext highlighter-rouge">external_experienced</code>)</li>
</ul>

<p>The site name matters: one company can run several sites, such as one for students and one for experienced hires. This parser pulls the three parts from any job or search link:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="nn">urllib.parse</span> <span class="kn">import</span> <span class="n">urlparse</span>

<span class="k">def</span> <span class="nf">parse_workday_link</span><span class="p">(</span><span class="n">link</span><span class="p">):</span>
    <span class="n">u</span> <span class="o">=</span> <span class="n">urlparse</span><span class="p">(</span><span class="n">link</span><span class="p">)</span>
    <span class="n">host</span> <span class="o">=</span> <span class="n">u</span><span class="p">.</span><span class="n">hostname</span> <span class="ow">or</span> <span class="s">""</span>
    <span class="k">if</span> <span class="ow">not</span> <span class="n">host</span><span class="p">.</span><span class="n">endswith</span><span class="p">(</span><span class="s">".myworkdayjobs.com"</span><span class="p">):</span>
        <span class="k">return</span> <span class="bp">None</span>
    <span class="n">tenant</span><span class="p">,</span> <span class="n">pod</span> <span class="o">=</span> <span class="n">host</span><span class="p">.</span><span class="n">split</span><span class="p">(</span><span class="s">"."</span><span class="p">)[:</span><span class="mi">2</span><span class="p">]</span>            <span class="c1"># "intel", "wd1"
</span>    <span class="n">parts</span> <span class="o">=</span> <span class="p">[</span><span class="n">p</span> <span class="k">for</span> <span class="n">p</span> <span class="ow">in</span> <span class="n">u</span><span class="p">.</span><span class="n">path</span><span class="p">.</span><span class="n">split</span><span class="p">(</span><span class="s">"/"</span><span class="p">)</span> <span class="k">if</span> <span class="n">p</span><span class="p">]</span>
    <span class="k">if</span> <span class="n">parts</span> <span class="ow">and</span> <span class="nb">len</span><span class="p">(</span><span class="n">parts</span><span class="p">[</span><span class="mi">0</span><span class="p">])</span> <span class="o">==</span> <span class="mi">5</span> <span class="ow">and</span> <span class="n">parts</span><span class="p">[</span><span class="mi">0</span><span class="p">][</span><span class="mi">2</span><span class="p">]</span> <span class="o">==</span> <span class="s">"-"</span><span class="p">:</span>
        <span class="n">parts</span> <span class="o">=</span> <span class="n">parts</span><span class="p">[</span><span class="mi">1</span><span class="p">:]</span>                        <span class="c1"># drop a locale such as "en-US"
</span>    <span class="n">site</span> <span class="o">=</span> <span class="n">parts</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="k">if</span> <span class="n">parts</span> <span class="k">else</span> <span class="bp">None</span>           <span class="c1"># "External"
</span>    <span class="k">return</span> <span class="p">{</span><span class="s">"tenant"</span><span class="p">:</span> <span class="n">tenant</span><span class="p">,</span> <span class="s">"pod"</span><span class="p">:</span> <span class="n">pod</span><span class="p">,</span> <span class="s">"site"</span><span class="p">:</span> <span class="n">site</span><span class="p">}</span>

<span class="k">print</span><span class="p">(</span><span class="n">parse_workday_link</span><span class="p">(</span>
    <span class="s">"https://adobe.wd5.myworkdayjobs.com/en-US/external_experienced/job/San-Jose/Senior-Engineer_R123"</span>
<span class="p">))</span>
<span class="c1"># {'tenant': 'adobe', 'pod': 'wd5', 'site': 'external_experienced'}
</span></code></pre></div></div>

<p>To find a company’s link, open any job on its careers page. If the address contains <code class="language-plaintext highlighter-rouge">myworkdayjobs.com</code>, you have it; otherwise check the apply button’s link.</p>

<h2 id="why-a-do-it-yourself-scraper-takes-longer-than-expected">Why a do-it-yourself scraper takes longer than expected</h2>

<p>Reading one site once is a weekend task. Keeping many sites running is not:</p>

<ul>
  <li><strong>The page is a JavaScript app.</strong> A plain HTTP request returns an empty shell, and the job list loads in pages.</li>
  <li><strong>Dates are text.</strong> Lists say “Posted 3 Days Ago” or “Posted 30+ Days Ago”, so older jobs have no exact date unless you open each job.</li>
  <li><strong>No salary field.</strong> Pay, when it is published at all, sits inside the description text.</li>
  <li><strong>Locations are shortened.</strong> A job in three cities shows as “US Washington DC (+2 more)”.</li>
  <li><strong>Large sites.</strong> Some employers list thousands of open jobs, and simple paging stops before the end.</li>
  <li><strong>Each employer sets its own rules.</strong> Check the site’s robots.txt before you read it, keep requests slow, and skip sites that do not allow automated access.</li>
</ul>

<p>None of this is impossible, but it repeats for every site you add and breaks quietly when a site changes.</p>

<h2 id="a-hosted-option-workday-jobs-api">A hosted option: Workday Jobs API</h2>

<p>I packaged that work as an Apify Actor, <a href="https://apify.com/conserving_celerytop/workday-jobs-api">Workday Jobs API</a>. You paste career site links (or company websites that link to Workday) and get one row per open job in a fixed format, read live, and only where the employer’s robots.txt allows it.</p>

<p>You need an Apify account and its API token (Console, Settings, API &amp; Integrations).</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"apify-client&gt;=3.2,&lt;4"</span>
<span class="nb">export </span><span class="nv">APIFY_TOKEN</span><span class="o">=</span>&lt;YOUR_APIFY_TOKEN&gt;
</code></pre></div></div>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">os</span>
<span class="kn">from</span> <span class="nn">decimal</span> <span class="kn">import</span> <span class="n">Decimal</span>
<span class="kn">from</span> <span class="nn">apify_client</span> <span class="kn">import</span> <span class="n">ApifyClient</span>

<span class="n">client</span> <span class="o">=</span> <span class="n">ApifyClient</span><span class="p">(</span><span class="n">os</span><span class="p">.</span><span class="n">environ</span><span class="p">[</span><span class="s">"APIFY_TOKEN"</span><span class="p">])</span>

<span class="n">run</span> <span class="o">=</span> <span class="n">client</span><span class="p">.</span><span class="n">actor</span><span class="p">(</span><span class="s">"conserving_celerytop/workday-jobs-api"</span><span class="p">).</span><span class="n">call</span><span class="p">(</span>
    <span class="n">run_input</span><span class="o">=</span><span class="p">{</span>
        <span class="s">"companies"</span><span class="p">:</span> <span class="p">[</span>
            <span class="s">"https://intel.wd1.myworkdayjobs.com/External"</span><span class="p">,</span>
            <span class="s">"https://adobe.wd5.myworkdayjobs.com/external_experienced"</span><span class="p">,</span>
        <span class="p">],</span>
        <span class="s">"jobFunctions"</span><span class="p">:</span> <span class="p">[</span><span class="s">"engineering"</span><span class="p">,</span> <span class="s">"data"</span><span class="p">],</span>
        <span class="s">"postedSince"</span><span class="p">:</span> <span class="s">"7 days"</span><span class="p">,</span>
    <span class="p">},</span>
    <span class="n">max_total_charge_usd</span><span class="o">=</span><span class="n">Decimal</span><span class="p">(</span><span class="s">"2.00"</span><span class="p">),</span>   <span class="c1"># a hard cap for this run
</span><span class="p">)</span>

<span class="k">for</span> <span class="n">job</span> <span class="ow">in</span> <span class="n">client</span><span class="p">.</span><span class="n">dataset</span><span class="p">(</span><span class="n">run</span><span class="p">.</span><span class="n">default_dataset_id</span><span class="p">).</span><span class="n">iterate_items</span><span class="p">():</span>
    <span class="k">if</span> <span class="n">job</span><span class="p">[</span><span class="s">"rowType"</span><span class="p">]</span> <span class="o">==</span> <span class="s">"job"</span><span class="p">:</span>
        <span class="k">print</span><span class="p">(</span><span class="n">job</span><span class="p">[</span><span class="s">"title"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">job</span><span class="p">[</span><span class="s">"location"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">job</span><span class="p">[</span><span class="s">"postedAt"</span><span class="p">],</span> <span class="s">"|"</span><span class="p">,</span> <span class="n">job</span><span class="p">[</span><span class="s">"url"</span><span class="p">])</span>
    <span class="k">else</span><span class="p">:</span>
        <span class="k">print</span><span class="p">(</span><span class="s">"status:"</span><span class="p">,</span> <span class="n">job</span><span class="p">[</span><span class="s">"company"</span><span class="p">],</span> <span class="n">job</span><span class="p">[</span><span class="s">"companyStatus"</span><span class="p">])</span>
</code></pre></div></div>

<p>Each job row has, among others:</p>

<table>
  <thead>
    <tr>
      <th>Field</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">title</code>, <code class="language-plaintext highlighter-rouge">department</code></td>
      <td>Job title and the category the site lists it under</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">location</code>, <code class="language-plaintext highlighter-rouge">countryCode</code>, <code class="language-plaintext highlighter-rouge">city</code></td>
      <td>Location text and parsed parts</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">workplaceType</code>, <code class="language-plaintext highlighter-rouge">remote</code></td>
      <td>Remote, hybrid or onsite</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">employmentType</code></td>
      <td>Workday’s time type, such as “Full time”</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">jobFunction</code>, <code class="language-plaintext highlighter-rouge">seniority</code></td>
      <td>Read from the title</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">postedAt</code></td>
      <td>Posted date; null for jobs 30 or more days old unless you add descriptions</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">url</code>, <code class="language-plaintext highlighter-rouge">applyUrl</code></td>
      <td>The job page</td>
    </tr>
  </tbody>
</table>

<p>Set <code class="language-plaintext highlighter-rouge">includeDescription</code> to <code class="language-plaintext highlighter-rouge">true</code> to add the description (with contact details removed), the tools it names, years of experience and any salary written in the text. A company with no matching jobs gets one status row that says why, such as <code class="language-plaintext highlighter-rouge">no_matching_jobs</code> or <code class="language-plaintext highlighter-rouge">not_found</code>.</p>

<h2 id="get-only-new-jobs-on-a-schedule">Get only new jobs on a schedule</h2>

<p>Add two fields and save the input as a task in Apify Console:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">run_input</span> <span class="o">=</span> <span class="p">{</span>
    <span class="s">"companies"</span><span class="p">:</span> <span class="p">[</span><span class="s">"https://intel.wd1.myworkdayjobs.com/External"</span><span class="p">],</span>
    <span class="s">"onlyNewJobs"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
    <span class="s">"monitorName"</span><span class="p">:</span> <span class="s">"workday-weekly"</span><span class="p">,</span>
    <span class="s">"alertWebhookUrl"</span><span class="p">:</span> <span class="s">"https://hooks.slack.com/services/..."</span><span class="p">,</span>   <span class="c1"># optional
</span><span class="p">}</span>
</code></pre></div></div>

<p>The first check returns the full list. Later checks return only jobs that are new or closed since the last one (<code class="language-plaintext highlighter-rouge">change</code> is <code class="language-plaintext highlighter-rouge">new</code> or <code class="language-plaintext highlighter-rouge">closed</code>) and can post an alert to Slack, Teams or any webhook. Schedule the task daily or weekly.</p>

<p>For one summary row per company, with open jobs, recent postings and counts per function, set <code class="language-plaintext highlighter-rouge">outputMode</code> to <code class="language-plaintext highlighter-rouge">companies</code>.</p>

<h2 id="what-it-costs">What it costs</h2>

<p>On September 27, 2026 the Actor charged $0.10 per company, with up to 1,000 jobs included, on Apify’s free plan ($0.09 on Scale, $0.07 on Business). Prices can change, so see the <a href="https://apify.com/conserving_celerytop/workday-jobs-api">Store page</a> for the current price. How it adds up:</p>

<ul>
  <li>$0.10 per company, also when a site lists no jobs or a website links no job board</li>
  <li>$0.05 for each further 1,000 jobs from one company</li>
  <li>$0.01 per started block of 200 descriptions, only when you ask for them</li>
  <li>$0.002 per later “only new jobs” check, per started 1,000 open jobs on the site</li>
</ul>

<p>So 100 companies cost $10 for the first read, and a daily check of the same 100 about $0.20. Sites whose robots.txt refuses access, invalid links and duplicates are free. <code class="language-plaintext highlighter-rouge">max_total_charge_usd</code> is a hard cap.</p>

<h2 id="limits-you-should-know">Limits you should know</h2>

<ul>
  <li><strong>It reads only Workday.</strong> For a list that mixes Greenhouse, Lever, Ashby and others, use <a href="https://apify.com/conserving_celerytop/live-career-page-jobs-api">ATS Jobs API</a>, which reads 22 job board systems including Workday, at $0.045 per company.</li>
  <li><strong>You bring the companies.</strong> It does not search all Workday employers. Paste links, websites, or pick a ready-made list.</li>
  <li><strong>Salary comes from the text.</strong> Workday has no pay field.</li>
</ul>

<h2 id="faq">FAQ</h2>

<p><strong>Is there an official public Workday jobs API?</strong>
Not for outsiders. Workday’s documented APIs need credentials from the employer. Public career sites can be viewed by anyone, and this Actor reads what they show.</p>

<p><strong>Can I get the recruiter’s name?</strong>
No. Rows hold job data only, and contact details are removed from descriptions.</p>

<hr />

<p><em>This article and the Actor are not affiliated with or endorsed by Workday, Inc. or any employer whose site it reads. Workday is a trademark of its owner.</em></p>]]></content><author><name>Don Mangu</name></author><category term="python" /><category term="api" /><category term="web-scraping" /><category term="apify" /><category term="recruitment" /><summary type="html"><![CDATA[Get open jobs from any Workday career site in Python: how the links work, why DIY scrapers break, and a hosted API at a flat price per company.]]></summary></entry></feed>