<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Salaheldinaz</title><link>https://salaheldinaz.com/</link><description>OSINT · Intelligence · Cyber Security · Automation · Technologist · Investigator</description><language>en-US</language><lastBuildDate>Sat, 10 Oct 2026 12:57:18 +0000</lastBuildDate><atom:link href="https://salaheldinaz.com/tags/automation/" rel="self" type="application/rss+xml"/><item><title>What automation actually means for an OSINT analyst</title><link>https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/</link><guid isPermaLink="true">https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/</guid><pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate><description>Not AI, not robots. Six stages that every collection system turns out to be a version of — and the 4 cases where the right answer is to do it by hand.</description><category>osint</category><category>automation</category><category>ai</category><category>methodology</category><content:encoded><![CDATA[<span class="hero-label not-prose">The Analyst&#39;s Stack · Foundations</span>

<p>I get asked some version of this question most weeks, usually by someone who is very
good at investigation and has decided they are bad at computers:</p>
<p><em>How do you actually use automation and AI in this work?</em></p>
<p>The honest answer is that the AI part is small and comes late. Almost all of the value
is in something much more boring, and the boring thing is learnable in an afternoon.
This series is the long answer. This post is the shape of it.</p>
<h2 id="first-what-it-is-not">First, what it is not</h2>
<p>Automation is not a robot doing your job. It is not a model that reads 1,000
documents and tells you which one matters. When people imagine automation and recoil,
they are usually imagining <strong>replacement</strong>, and then quite reasonably concluding that
their judgement cannot be replaced. They are right. That is not what this is.</p>
<p>Automation is also not &ldquo;being technical.&rdquo; The most useful automated thing I have ever
built was 40 lines long and did one thing: it checked a website every 15 minutes
and told me when something new appeared. It made no decisions. It had no opinions. It
just meant that I found out in 15 minutes instead of the following Tuesday, and
that turned out to be the entire difference between a story and a missed story.</p>
<p>Here is a better definition:</p>
<div class="callout callout-info not-prose"><p class="callout-title">Working definition</p><div class="callout-body">Automation is <strong>moving the act of looking from your attention to a machine&rsquo;s</strong>, so that
your attention is spent on the part only you can do — deciding what a thing means.</div>
</div>

<p>That is it. Everything else in this series is detail.</p>
<h2 id="every-collection-system-is-the-same-6-stages">Every collection system is the same 6 stages</h2>
<p>Once you have built a few of these, you notice something slightly deflating: they are
all the same system. The subject changes, the sources change, the language changes, and
the shape does not. Every collection system I have built, and every one I have taken
apart, is some version of 6 stages in a row. I call that shape <strong>the spine</strong>, and I
will keep calling it that for the rest of the series.</p>
<figure class="diagram not-prose">
  <img class="diagram-light" src="https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/images/f1-spine_light.svg" width="1160" height="214" alt="Six stages in a row, left to right: Collect, Normalize, Store, Analyze, Alert, Present. Analyze is highlighted." loading="lazy" decoding="async">
  <img class="diagram-dark" src="https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/images/f1-spine_dark.svg" width="1160" height="214" alt="Six stages in a row, left to right: Collect, Normalize, Store, Analyze, Alert, Present. Analyze is highlighted." loading="lazy" decoding="async">
  <figcaption class="post-caption"><strong>The spine.</strong> Every post in this series refers back to this. Plenty of systems skip a stage. Almost none reorder it, and part 9 says so on its page.</figcaption>
</figure>

<p>Read left to right:</p>
<div class="step not-prose">
  <div class="step-head">
    <span class="step-num">1</span>
    <span class="step-title">Collect</span>
  </div>
  <div class="step-body">Get the raw thing. A page, a post, a file, an image, a row from an API <em>(a structured
way for one program to ask another for data)</em>. This stage is where most of the fighting
happens — sites that block you, sources with no export button, formats that change
without warning — but conceptually it is the simplest step in the chain.</div>
</div>

<div class="step not-prose">
  <div class="step-head">
    <span class="step-num">2</span>
    <span class="step-title">Normalize</span>
  </div>
  <div class="step-body">Turn the raw thing into one consistent shape. A post from one platform and a post from
another arrive looking nothing alike; after this stage they both have a date, an author,
a body, a link, and a source. Everything downstream depends on this, and almost everyone
skips it at first.</div>
</div>

<div class="step not-prose">
  <div class="step-head">
    <span class="step-num">3</span>
    <span class="step-title">Store</span>
  </div>
  <div class="step-body">Write it down somewhere that survives the program exiting. Usually a small database. The
critical part is not the storage — it is giving every item a <strong>stable identifier</strong> so
that you can tell, tomorrow, whether you have already seen it.</div>
</div>

<div class="step not-prose">
  <div class="step-head">
    <span class="step-num">4</span>
    <span class="step-title">Analyze</span>
  </div>
  <div class="step-body">Narrow the pile. Translate, classify, extract names and places, group near-duplicates,
score for relevance. This is the only stage where AI genuinely belongs, and even here it
is narrowing, never deciding. Two later posts — <span class="link-pending" title="Coming soon in this series">where AI actually fits in
OSINT</span> and <span class="link-pending" title="Coming soon in this series">the do&rsquo;s and don&rsquo;ts for
investigators</span> — are about this stage and
nothing else.</div>
</div>

<div class="step not-prose">
  <div class="step-head">
    <span class="step-num">5</span>
    <span class="step-title">Alert</span>
  </div>
  <div class="step-body">Tell a human that something crossed a threshold. A message, an email, a row turning red.
If nothing ever reaches a person, you have built a very expensive way of generating logs.</div>
</div>

<div class="step not-prose">
  <div class="step-head">
    <span class="step-num">6</span>
    <span class="step-title">Present</span>
  </div>
  <div class="step-body">Make it readable. A map, a table, a timeline, a feed. This is where an investigation
actually happens, and it is the stage most often treated as an afterthought.</div>
</div>

<p>You will see this diagram again on nearly every post in this series, because the fastest
way to understand an unfamiliar tool is to work out which stages it covers and which it
leaves to you.</p>
<h2 id="the-2-stages-nobody-wants-to-build">The 2 stages nobody wants to build</h2>
<p>Ask someone to describe the system they want and they will describe stages 1, 5
and 6: <em>get the posts, ping me, show me a map.</em> Nobody has ever asked me for stage 2
or 3. They are also the 2 stages that decide whether the thing survives contact with
a month of real use.</p>
<p>Here is why, in the most concrete way I can put it.</p>
<p>Say you are watching a site that posts new entries. You write something that fetches the
page every 15 minutes and messages you about each entry it finds. You run it. It
works. You go to bed.</p>
<p>By morning you have 96 identical messages about the same entry, because &ldquo;the
entries currently on the page&rdquo; is not the same question as &ldquo;the entries I have not seen
before.&rdquo; The program has no memory. Every run is its first run.</p>
<p>The fix is stage 3, and it is smaller than you would guess: give each entry an
identifier that does not change, write it down, and skip anything you have already
written down. Now running the program twice produces nothing the second time. That
property has a name — <strong>idempotency</strong> <em>(running the same job twice leaves you in the
same state as running it once)</em> — and once you start looking for it you will notice it
is the difference between tools that people keep using and tools that get switched off
in week 2.</p>
<div class="callout callout-default not-prose"><p class="callout-title">The cheapest possible version</p><div class="callout-body">The identifier does not need to be clever. A hash of the entry&rsquo;s URL plus its title is
usually enough — a hash <em>(a short fixed-length fingerprint of some text; the same input
always produces the same fingerprint, and a different input almost never does)</em> gives
you a stable key even when the source gives you nothing to work with. Store those keys
in a file if you like. It does not have to be a database to be a database.</div>
</div>

<figure class="diagram not-prose">
  <img class="diagram-light" src="https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/images/f1-hash_light.svg" width="1000" height="268" alt="Two entries from the same page. The first entry&#39;s URL and title are joined and hashed into a short key. Running the job again on the same entry produces the same key, so it is skipped. A second, different entry produces a completely different key and is treated as new." loading="lazy" decoding="async">
  <img class="diagram-dark" src="https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/images/f1-hash_dark.svg" width="1000" height="268" alt="Two entries from the same page. The first entry&#39;s URL and title are joined and hashed into a short key. Running the job again on the same entry produces the same key, so it is skipped. A second, different entry produces a completely different key and is treated as new." loading="lazy" decoding="async">
  <figcaption class="post-caption"><strong>The same entry always lands on the same key.</strong> Both keys here are real: run SHA-256 over the string on the left and the string in the middle is the first 12 characters of the result.</figcaption>
</figure>

<p>Stage 2 earns its keep the moment you add a second source. As long as you have one
source, normalizing feels like pointless ceremony. Add a second and you either normalize
or you write every downstream piece of logic twice. Add a fifth and the version without
a normalize stage is no longer repairable.</p>
<p>Naming the stages costs nothing and buys you one thing: a run that tells you where it got
to. Here is a toy version of the whole spine — 2 invented feeds that arrive in different
shapes, one line printed as each stage finishes.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">$ python3 spine.py
</span></span><span class="line"><span class="cl">collect    2 sources polled, 7 raw items
</span></span><span class="line"><span class="cl">normalize  7 items in one shape: source, title, link, date, author
</span></span><span class="line"><span class="cl">store      7 seen, 7 new, 0 already known
</span></span><span class="line"><span class="cl">analyze    7 new items narrowed to 3 worth a look
</span></span><span class="line"><span class="cl">alert      3 alerts written to alerts.txt
</span></span><span class="line"><span class="cl">present    wrote report.html
</span></span></code></pre></div><div class="post-caption not-prose"><strong>One run, printing as it goes.</strong> The value of naming the stages is that a failure tells
you which one broke.</div>

<p>Run it a second time and the <code>store</code> line reads <em>7 seen, 0 new, 7 already known</em>, and the
analyze and alert lines fall to zero. That is idempotency, printing itself.</p>
<h2 id="where-ai-actually-comes-in">Where AI actually comes in</h2>
<p>Stage 4, and only stage 4.</p>
<p>I will spend 2 whole posts on this later, so here I will just plant the flag. In this
work, a model is good at reducing a pile you cannot read into a pile you can. Translate
400 posts so you can skim them. Sort 10,000 items into &ldquo;probably
relevant&rdquo; and &ldquo;probably not.&rdquo; Notice that these 6 records are describing the same
event under 6 spellings.</p>
<p>It is not good at telling you whether something is true. It can attempt verification and
attribution, but what comes back is a lead, not a finding, and checking it is your job.
The moment a model&rsquo;s output becomes evidence without a human between it and the
conclusion, you have automated the one part of the job that was never supposed to be
automated.</p>
<p>AI narrows the pile. A person decides what is in it. If you cannot point at the human
who checked a given claim, the claim is not checked, and &ldquo;the model was confident&rdquo; is not
a sourcing note you can publish.</p>
<div class="limitation not-prose"><p class="limitation-title">What the spine does not tell you</p><div class="limitation-body"><p>It is a shape, not a plan, and it is silent on the 3 things most likely to sink the
project.</p>
<p><strong>It says nothing about whether you are allowed to.</strong> Six stages will happily collect
something you should not be holding. The legal and ethical question is not a stage; it
sits before stage 1 and the diagram cannot remind you of it.</p>
<p><strong>It says nothing about cost.</strong> A pipeline with all 6 stages can run for pennies or can
quietly spend more on model calls in a weekend than the story is worth. Nothing in the
shape tells you which.</p>
<p><strong>A complete system can still be useless.</strong> I have built things with all 6 stages
working perfectly that nobody opened twice, because the question they answered was one I
had invented. The spine tells you whether a tool is finished. It cannot tell you whether
it was worth building.</p>
</div>
</div>

<h2 id="when-not-to-automate">When not to automate</h2>
<p>This is the part that gets left out of tutorials, because tutorials are written by
people who want you to finish the tutorial.</p>
<figure class="diagram not-prose">
  <img class="diagram-light" src="https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/images/f1-when-not_light.svg" width="1000" height="286" alt="Two rows of conditions. Automate when it repeats, the source holds its shape, the volume is past eyeballing, and you would notice it stopping. Do not automate when the question is asked once, the volume is a handful, the source changes shape weekly, or the legal ground is unsettled." loading="lazy" decoding="async">
  <img class="diagram-dark" src="https://salaheldinaz.com/blog/what-automation-means-for-an-osint-analyst/images/f1-when-not_dark.svg" width="1000" height="286" alt="Two rows of conditions. Automate when it repeats, the source holds its shape, the volume is past eyeballing, and you would notice it stopping. Do not automate when the question is asked once, the volume is a handful, the source changes shape weekly, or the legal ground is unsettled." loading="lazy" decoding="async">
  <figcaption class="post-caption"><strong>Four reasons to put the laptop down.</strong> Any one of the bottom row is enough.</figcaption>
</figure>

<p><strong>The question gets asked once.</strong> Automation is an investment that pays back over
repetitions. If you need this answer today and never again, open the site and read it.
I have watched people — I have <em>been</em> people — spend 6 hours building something to
avoid 90 minutes of clicking.</p>
<p><strong>The volume is small.</strong> Forty items is not a data problem. Forty items is an afternoon.
The threshold where a machine beats a careful human is higher than it feels when you are
bored.</p>
<p><strong>The source changes shape every week.</strong> Some sites are stable for years. Others rewrite
their front end constantly, and every rewrite breaks your collection silently — which is
worse than breaking loudly, because a tool that quietly returns nothing looks exactly
like a week with no news. If you cannot detect that failure, do not build on that source.</p>
<p><strong>The legal or ethical ground is unsettled.</strong> Terms of service, the rules that apply to
personal data where you work, and the question of whether collecting at scale changes
what a thing <em>is</em> — all of that gets decided before the first line of code, not after
someone asks you about it. This is not a technical constraint and you cannot engineer
around it.</p>
<p>Of the 4, 2 are about restraint and the other 2 about honesty, which is roughly
the ratio I would expect.</p>
<h2 id="what-is-in-the-rest-of-this-series">What is in the rest of this series</h2>
<p>There are 22 posts, of 2 kinds that work differently.</p>
<p><strong>The guide</strong> is 10 teaching posts. Four more foundations follow this one: building a
first workflow end to end, where AI fits, the do&rsquo;s and don&rsquo;ts, and the OPSEC problem
that automation creates <em>(you have automated the act of looking, so you have automated
the act of being seen)</em>. Then 5 methods that go deeper — automating without writing
code, running models on your own machine, reaching sources that block you, measuring
whether an AI step actually works, and storing what you collect so you can find it
again.</p>
<p><strong>The showcase</strong> is 12 short posts, one per working system. Why it was built, what
it solves, how it fits together, and what it cannot do.</p>
<p>If you read only two, read the next one and the one on measuring AI steps. The first
gets something running. The second is the one that will stop you publishing something
wrong.</p>
<h2 id="start-here">Start here</h2>
<p>If you want something concrete before the next post lands: pick one source you check by
hand more than twice a week. Not your most important one. Your most repetitive one.
Write down, in plain words, what you do each time you check it, one line per step.</p>
<p>You have just written most of the spine. The rest is typing.</p>
<p>If you have never written Python and want to run the code in the <span class="link-pending" title="Coming soon in this series">next
post</span>, it lists 3 free courses in the
order I would take them: futurecoder from a standing start, PyBasic for practice, then
Python for OSINT in 21 days for the same skills aimed at investigation work.</p>
<aside class="related-earlier not-prose" aria-labelledby="related-earlier-heading">
  <h2 id="related-earlier-heading" class="related-earlier-title">Related earlier work</h2>
  <p class="related-earlier-intro">Outside the series, but each one is a worked example of a pattern it argues for.</p>
  <ul class="related-earlier-list">
    <li class="related-earlier-item">
      <a href="https://salaheldinaz.com/blog/pwtt-qgis-plugin/">Track the Unseen - Spot Change Anywhere, Any Time</a>
      <span class="related-earlier-note">A desktop tool wrapping someone else&#39;s analysis method, so the method survives past the paper it was published in.</span>
    </li>
    <li class="related-earlier-item">
      <a href="https://salaheldinaz.com/blog/google-map-exporter/">Google Map Exporter — Export Any Google Map to KML, KMZ, or GeoJSON</a>
      <span class="related-earlier-note">Automating a source that publishes no API and does not want to be read in bulk.</span>
    </li>
    <li class="related-earlier-item">
      <a href="https://salaheldinaz.com/blog/adsb-history-tool/">Adsb-History Self-Hosted</a>
      <span class="related-earlier-note">Turning a live feed into a queryable history — the collect-then-store half of the pipeline, on its own.</span>
    </li>
    <li class="related-earlier-item">
      <a href="https://salaheldinaz.com/blog/wigle-to-google-earth/">Wigle to Google Earth</a>
      <span class="related-earlier-note">A format-conversion step small enough to be worth ten minutes of scripting and not one more.</span>
    </li>
  </ul>
</aside>
]]></content:encoded></item></channel></rss>