<?xml version="1.0" encoding="UTF-8" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Blogit</title>
    <link>https://julianmoors.net/</link>
    <description>A secure, server-rendered blog application</description>
    <language>en-us</language>
    <lastBuildDate>Sat, 12 Sep 2026 05:25:23 GMT</lastBuildDate>
    <atom:link href="https://julianmoors.net/rss" rel="self" type="application/rss+xml" />
    
    <item>
      <title>How to Build a Knowledge Base Your AI Agents Will Love</title>
      <link>https://julianmoors.net/posts/how-to-build-a-knowledge-base-your-ai-agents-will-love</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/how-to-build-a-knowledge-base-your-ai-agents-will-love</guid>
      <pubDate>Sat, 05 Sep 2026 19:58:05 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>AI agents are only as good as the information they can access. If you want your agents to answer questions, complete tasks, or help customers accurately, they need a well-organised knowledge base to draw from. Here is how to build one that actually works.</p>
<p>Start with a clear structure. A knowledge base is not a random pile of documents. It should be organised into logical sections — policies, product information, how-to guides, and so on — so both humans and AI can find things easily. Good structure makes everything downstream easier.<br />Write in plain, specific language. AI agents understand clear writing better than vague prose. Avoid jargon where possible, define terms, and write in short, focused paragraphs. If a human would find a document confusing, an AI agent probably will too.</p>
<p>Keep each document focused on one topic. A page that tries to cover ten different things is hard to navigate and hard for an AI to use. Break things up into small, single-purpose documents that answer one question well.</p>
<p>Keep it up to date. The most common reason AI agents give wrong answers is that the knowledge base is stale. Set up a process for reviewing and updating content regularly, and assign someone to own it. Outdated information is worse than no information.</p>
<p>Make it searchable and accessible. The knowledge base needs to be somewhere your AI agents can actually read, whether that is a shared drive, a wiki, or files in a repository. The exact tool matters less than making sure the agent can reliably pull from it.</p>
<p>Finally, test it. Ask your AI agent real questions and see how well it answers. When it gets something wrong, fix the source document, not just the answer. Over time, that feedback loop turns your knowledge base into a genuinely reliable foundation for your agents.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1788638280679.jpeg" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>The Gated Mind: Are Frontier AI Labs Overcharging Themselves into Oblivion?</title>
      <link>https://julianmoors.net/posts/the-gated-mind-are-frontier-ai-labs-overcharging-themselves-into-oblivion</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/the-gated-mind-are-frontier-ai-labs-overcharging-themselves-into-oblivion</guid>
      <pubDate>Fri, 28 Aug 2026 16:22:32 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p><strong>Frontier AI companies risk pricing themselves out of democratic utility by building walled corporate gardens, making advanced intelligence an exclusive asset for massive corporations.</strong> As token consumption models scale toward autonomous agents, many enterprises are finding that running closed frontier models costs more than the human workflows they were meant to optimize. This economic friction raises a critical question for the industry: Are Silicon Valley’s elite labs shooting themselves in the foot by gatekeeping the foundational cognitive infrastructure of the future?</p>
<h2>The Metered Brain: A Utility for the Few?</h2>
<p>The current trajectory of closed AI models treats intelligence not as an open infrastructure, but as a metered commodity. During an appearance at the BlackRock Infrastructure Summit, OpenAI CEO Sam Altman outlined this corporate vision cleanly:</p>
<blockquote>
<p>"We see a future where intelligence is a utility, like electricity or water, and people buy it from us on a meter, and use it for whatever they want to use it for."</p>
</blockquote>
<p>While Altman frequently argues that the price of intelligence drops 10x every year, his metered framework introduces a structural risk. If a handful of centralized tech giants control the infrastructure and gatekeep access via proprietary APIs, raw intelligence becomes a luxury asset.</p>
<p>When computing infrastructure lags behind global demand, Altman admits that prices could spike, forcing advanced AI into the hands of the wealthy or under the direction of bureaucratic central planning. This economic model threatens to lock out smaller startups, researchers, and public institutions, leaving only multi-billion-dollar corporations with the capital to pay the "intelligence bill".</p>
<h3>The Sovereign Data Crisis</h3>
<p>Beyond the pure balance sheet, closed-source models create a massive security flaw for global institutions: the surrender of data sovereignty.</p>
<p>When a nation, a health network, or a defense contractor routes its internal knowledge bases through a centralized corporate API, they lose absolute control over that data. For national governments and highly regulated industries, trusting an external, US-centric corporate black box with critical infrastructure analytics is a non-starter. Walled gardens require organizations to hand over their unique intellectual property just to query the system, turning localized knowledge into training material for future generic corporate models.</p>
<p><strong>Open Weights: The Democratic Counter-Weight</strong></p>
<p>Fortunately, the market is aggressively correcting. Open-weight models are quickly closing the capability gap with closed corporate systems, turning raw intelligence into a true commodity.</p>
<p>Deploying highly capable open-weight models locally offers several massive advantages over centralized APIs:</p>
<p><strong>Absolute Data Sovereignty</strong>: Highly sensitive datasets remain fully on-premise, safely contained behind local firewalls without zero-day corporate sandbox leaks.</p>
<p><strong>Zero Token Cost Friction</strong>: Organizations bypass volatile, usage-based subscription tiers by running models natively on localized hardware.</p>
<p><strong>Granular Customization</strong>: Developers can modify underlying architecture and fine-tune parameters without corporate content moderation filters or arbitrary platform updates.</p>
<p>By treating intelligence as open-source code rather than a metered utility, societies can build localized, resilient digital infrastructure. If frontier labs continue to demand high margins for raw reasoning tokens, they won't just alienate developers—they will force the rest of the world to build an alternative ecosystem entirely out of their reach.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1787934626794.jpg" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>The Scaffolding Revolution: Why DeepSeek’s New Harness is Breaking the Internet</title>
      <link>https://julianmoors.net/posts/the-scaffolding-revolution-why-deepseek-s-new-harness-is-breaking-the-internet</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/the-scaffolding-revolution-why-deepseek-s-new-harness-is-breaking-the-internet</guid>
      <pubDate>Sun, 23 Aug 2026 09:42:40 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<h3>DeepSeek Harness has officially shattered records, blowing past 175,000 GitHub stars in its very first week of release.</h3>
<p>This represents one of the fastest adoption curves the developer community has ever seen. While most people get excited about the actual artificial intelligence "brains," DeepSeek just proved that the "body"—how the AI connects to your computer and actually does work—is what everyone has been waiting for.</p>
<h2>What Exactly is an AI "Harness"?</h2>
<p>Think of a standard AI model like a brilliant chef who is locked inside a room with no doors, no windows, and no utensils. The chef can give you an incredible recipe, but they cannot actually cook the meal for you.</p>
<p>An AI harness is the kitchen. It gives the AI:</p>
<ul>
<li><p>A digital hands-and-eyes system to look at your files</p>
</li>
<li><p>A terminal to run software commands</p>
</li>
<li><p>The ability to coordinate other AI helpers to get a massive project done</p>
</li>
</ul>
<p>Until now, the coolest kitchens belonged to closed systems like Anthropic's Claude Code or SpaceX-backed Cursor. They are incredibly powerful, but you are forced to use their rules, their models, and their pricing. DeepSeek changed the game by releasing their entire agent framework under a fully free, open-source MIT license.</p>
<h2>The "Everything is a Plugin" Superpower</h2>
<p>The real reason tech communities are losing their minds over DeepSeek Harness is its unique architectural philosophy: <strong>Everything is a plugin</strong>.</p>
<p>In other AI setups, changing how the AI acts requires rewriting thousands of lines of complex internal code. With DeepSeek Harness, the whole system is built out of Lego blocks.</p>
<ul>
<li><p><strong>Swap the AI Brain:</strong> Don't want to use DeepSeek's model? You can use a YAML configuration file to point it at OpenAI, Anthropic, or a completely free model running directly on your laptop.</p>
</li>
<li><p><strong>Mix and Match:</strong> You can actually wire competitor tools like Claude Code right into the harness as a plugin. DeepSeek can act as the manager who delegates specific coding tasks to Claude.</p>
</li>
<li><p><strong>Self-Improving Agents:</strong> The system is so flexible that the AI can actually write and install new plugins for itself while it's running.</p>
</li>
</ul>
<h2>Getting It Running Locally</h2>
<p>Unlike other developer tools that force you to stay locked inside a boring command terminal, DeepSeek ships with a clean, local Web UI out of the box.</p>
<p>Setting it up takes less than a minute. If you have Node.js installed, you just open your computer terminal and type a single command:</p>
<div class="codeblock-wrapper bg-gray-50 border border-gray-200 rounded-lg my-6"><pre class="rounded-lg py-3 px-4 my-0 overflow-x-auto language-javascript"><code class="language-javascript">npx @deepseek-ai/dsh web
</code></pre></div>

<p>This instantly boots up a sleek visual dashboard right in your browser at <code>http://127.0.0.1:3080</code>. From there, you get a front-row seat to how your AI is thinking, including real-time token speeds, tool histories, and cost tracking.</p>
<h2>A Quick Reality Check</h2>
<p>Before you go completely replacing your entire workflow with this new tool, remember that DeepSeek Harness is currently in developer preview (version 0.1).</p>
<p>The creators have openly stated in their official DeepSeek GitHub README that they are iterating rapidly, and future updates will break existing configurations. It has some rough edges, the documentation can be a bit confusing, and it can eat through data quickly if you aren't careful.</p>
<p>Think of it as a thrilling, high-performance concept car. It is the absolute best way to see where the future of software development is heading, but you might want to use it for side experiments before trusting it with your most important production data.</p>
<p>The era of closed-off, rigid AI workflows is officially ending, and DeepSeek just handed the keys of customization back to the open-source community.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1787478524114.webp" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>Announcing Aspect Themes - Premium Admin Dashboards</title>
      <link>https://julianmoors.net/posts/announcing-aspect-themes-premium-admin-dashboards</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/announcing-aspect-themes-premium-admin-dashboards</guid>
      <pubDate>Sat, 15 Aug 2026 05:56:43 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>I am thrilled to announce the official launch of my brand new website, Aspect Themes!</p>
<p>After months of hard work, design iterations, and deep-dive development, this platform is finally live to change how you build and interact with user interfaces. If you have ever stared at a clunky admin panel and thought "there has to be a better way," this one is for you.</p>
<h2>Why Aspect Themes?</h2>
<p>For a long time, I have noticed a massive gap between beautiful website interfaces and functional backend admin areas. Most admin dashboards are either incredibly cluttered or painfully boring to look at. The front end gets all the love, while the tools your team actually uses every day get left behind.</p>
<p>I built Aspect Themes to solve exactly that. It is a dedicated hub for beautifully crafted, highly functional, and fully responsive <strong>admin dashboard templates</strong>. These themes are built specifically for developers, creators, and businesses who want their behind-the-scenes systems to feel just as premium as their public-facing websites. In plain terms, your admin area should be a joy to use, not a chore.</p>
<h2>Built for the Modern Stack</h2>
<p>As a developer, I know how frustrating it is to work with bloated, heavy codebases. That is why I built these themes using the tools we actually love to use:</p>
<p><strong>Tailwind CSS Utility-First Design</strong>: No messy CSS overrides. Every component is built natively with Tailwind, making customisation incredibly fast and straightforward.</p>
<p><strong>Seamless Node.js Integration</strong>: The architecture is fully optimized for Node.js environments, ensuring rapid rendering and smooth backend connectivity.</p>
<p><strong>Native Light &amp; Dark Modes</strong>: No lazy color inversions here. Every theme includes meticulously designed light and dark modes that adapt beautifully to user preferences with a single toggle.</p>
<p><strong>Data-Driven Focus</strong>: Deeply optimized layouts for complex charts, tables, user management metrics, and comprehensive analytics tracking.</p>
<h2>This Is Just the Beginning!</h2>
<p>Launching this project is a huge milestone for me, but the real work starts now. I will be continuously adding new template variations, custom components, and specialized framework integrations to the catalog. I have a long list of ideas queued up, and I can't wait to share each one with you.</p>
<p>Go ahead and check out the live marketplace at <a href="https://www.aspectthemes.com" target="_blank" rel="noopener noreferrer">Aspect Themes</a>. I would absolutely love to hear your thoughts, feedback, or any feature requests you might have!</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1788368068553.png" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>Unleashing the ultimate source of truth: OpenCode, Obsidian, and Agentic Development</title>
      <link>https://julianmoors.net/posts/unleashing-the-ultimate-source-of-truth-opencode-obsidian-and-agentic-development</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/unleashing-the-ultimate-source-of-truth-opencode-obsidian-and-agentic-development</guid>
      <pubDate>Tue, 21 Jul 2026 18:27:21 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>Managing a project's "Source of Truth" (SoT) is a historical developer headache. Documentation sits in corporate wikis, requirements live in Jira, and code lives in Git. They constantly drift apart.</p>
<p>By pairing OpenCode (an alternative to Claude Code and Codex) with Obsidian (a local-first, Markdown-based knowledge base), you can build a unified, human-readable, and machine-actionable SoT.</p>
<p>When you layer Agentic Driven Development (ADD) on top of this stack, the documentation stops being a passive diary and becomes an active participant in building your software.</p>
<h2>The Stack: Git + OpenCode + Obsidian</h2>
<p>Using proprietary, siloed tools for project management creates platform lock-in and fragmented context. Here is why the open, local-first combination wins:</p>
<ul>
<li><p>Everything is Markdown: Obsidian stores files as plain Markdown documents. OpenCode projects produce text files. Because everything is text, your code and your documentation can live in the exact same Git repository.</p>
</li>
<li><p>Bi-Directional Linking: Obsidian’s core strength is linking notes together. You can link a feature request to a specific technical design document, which links directly to local code snippets.</p>
</li>
<li><p>Version-Controlled Truth: When documentation lives inside your project repository, changes to documentation are tracked via Git commits alongside the code. A feature pull request updates the code and the documentation simultaneously.</p>
</li>
</ul>
<h2>Why Agentic Driven Development is the Perfect Match</h2>
<p>While Behavior-Driven Development (BDD) is excellent for turning user stories into automated tests, Agentic Driven Development (ADD) is far more appropriate for this specific stack.</p>
<p>ADD views AI agents as autonomous collaborators rather than just autocomplete tools. Agents require structured, highly contextual, and easily readable data to operate effectively. An Obsidian vault hosted on an open repository acts as the perfect cognitive map for an AI agent.</p>
<ol>
<li>The Context Engine for AI Agents</li>
</ol>
<p>AI agents struggle with massive, disorganized codebases. By structuring your project architecture, business logic, and roadmap inside Obsidian using clear Markdown hierarchies and tags, you create a map the agent can instantly scan. The agent reads your Obsidian notes to understand why a feature is being built before it touches the code.</p>
<ol start="2">
<li>Autonomous Documentation Sync</li>
</ol>
<p>In an ADD workflow, the agent doesn't just write code; it maintains the SoT. When an agent refactors a backend service, it can autonomously update the corresponding architectural diagrams (using Mermaid.js inside Obsidian) and API endpoints documented in your vault.</p>
<ol start="3">
<li>Executable Roadmaps</li>
</ol>
<p>With your project source code and Obsidian, your project roadmap is just a text file. An AI agent can parse your Roadmap.md or Todo.md, identify the next highest-priority task, spin up a new git branch, write the code, test it, and submit a pull request—all while updating the task status in your vault</p>
<h2>How to Set It Up</h2>
<ol>
<li><p>Initialize a Unified Repo: Create an open-source repository. Inside it, create a /src folder for your code and a /docs folder for your Obsidian vault.</p>
</li>
<li><p>Link the Vault: Open the /docs folder as a new vault in Obsidian.</p>
</li>
<li><p>Define Agent Prompts: Instruct your development agents (like Cline, Roo Code, or custom LLM scripts) to always check the /docs folder first for context before writing code, and to update those docs when a feature changes.</p>
</li>
</ol>
<h2>The Verdict</h2>
<p>When your code is open, your documentation is a local-first network of Markdown files, and your primary developer is an AI agent, you eliminate project friction. You no longer write documentation for a dusty corporate wiki; you write it to program your AI agents.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1785666182331.jpg" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>Why companies are swapping SaaS for AI</title>
      <link>https://julianmoors.net/posts/why-companies-are-swapping-saas-for-ai</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/why-companies-are-swapping-saas-for-ai</guid>
      <pubDate>Tue, 21 Jul 2026 12:49:54 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>Traditional SaaS companies charge you per user, per month. As your team grows, your software bill skyrockets. It is one of those quiet costs that creeps up on you — you add three new hires and suddenly you are paying for three more seats across half a dozen tools.</p>
<p>With advanced, low-cost AI models like DeepSeek, a company can build its own basic customer service bots, data analysis tools, and writing assistants. Instead of buying ten different software tools, you route your workflows through one AI system and pay pennies for the actual computing power used. That is a fundamentally different way to think about software.</p>
<h2>The Cost Breakdown: SaaS vs. AI</h2>
<p>Let's look at a realistic scenario for a mid-sized company with 50 employees using standard tools for copywriting, customer support, and CRM data management.</p>
<h3>1. The Traditional SaaS Route</h3>
<ul>
<li>Average SaaS cost: £40 per user, per month (across 3 essential tools)</li>
<li>Monthly bill: £2,000</li>
<li>6-Month Cost: £12,000</li>
<li>12-Month Cost: £24,000</li>
</ul>
<p>That £24,000 a year buys you convenience, but it also buys you a lot of unused seats, overlapping features, and rising prices every renewal cycle.</p>
<h3>2. The DeepSeek AI API Route</h3>
<p>Using DeepSeek's highly efficient API models, you only pay for the volume of text generated (measured in tokens). For a team of 50 heavily using AI daily, they might process 100 million tokens a month.</p>
<ul>
<li>Average AI data cost: £100 per month</li>
<li>Internal setup &amp; hosting: £150 per month</li>
<li>Total monthly bill: £250</li>
<li>6-Month Cost: £1,500</li>
<li>12-Month Cost: £3,000</li>
</ul>
<p>The difference is stark: £3,000 instead of £24,000 for the same underlying needs.</p>
<h3>The Catch: Upfront Investment</h3>
<p>While the operational savings are massive, you must consider the initial setup. You will need to hire a developer or use an agency to build the internal interface for your staff.</p>
<p>If building a custom dashboard costs your company £5,000 upfront:</p>
<ul>
<li>Your 6-month savings adjust to £5,500</li>
<li>Your 12-month savings jump to £16,000</li>
</ul>
<p>From Year 2 onwards, the software is completely yours, and the ongoing savings go straight to your bottom line. You also avoid the annual price hikes and feature-lock that come with traditional SaaS subscriptions.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1785666149174.jpg" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>How I slashed 90% of API costs using this one AI model</title>
      <link>https://julianmoors.net/posts/how-i-slashed-90-of-api-costs-using-this-one-ai-model</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/how-i-slashed-90-of-api-costs-using-this-one-ai-model</guid>
      <pubDate>Mon, 15 Jun 2026 18:28:19 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>You can probably tell from this blog post's cover image which AI model I'm talking about. And yes, DeepSeek v4 Pro is ~90% cheaper than the other frontier models. How does DeepSeek achieve this?</p>
<p>Here's what you need to know:</p>
<h3>1. Mixture-of-Experts (MoE)</h3>
<p>DeepSeek uses Mixture-of-Experts to balance model capacity with computational efficiency.</p>
<ul>
<li>Specialized Experts: Each feedforward block is divided into multiple "experts" (e.g., V3: 256 experts). Instead of every token passing through all parameters, only a few top-ranked experts (2–4 per token) are activated by a routing mechanism that scores token embeddings against expert embeddings.</li>
<li>Routing Mechanism: A sigmoid gating function with learnable expert biases dynamically selects which experts are engaged for each token. This resolves load-imbalance problems common in MoE architectures without adding auxiliary-loss terms.</li>
<li>Compute Savings: MoE allows DeepSeek to behave like a 671B parameter dense model on training while activating only ~5.5% (37B) of parameters during inference. This reduces both GPU memory usage and runtime costs, enabling long-context capabilities without prohibitive expenditure.</li>
<li>Shared Experts &amp; Redundancy: Certain popular experts can be duplicated across GPUs in deployment to avoid bottlenecks. This ensures throughput remains high under uneven token distribution.</li>
</ul>
<h3>2. KV Cache and Multi-Head Latent Attention (MLA)</h3>
<p>Transformer models rely on storing Key (K) and Value (V) tensors for each token to compute attention efficiently over long contexts—this is the KV cache. For extremely long sequences (e.g., 128K tokens in DeepSeek), naive caching would cause memory usage to skyrocket.</p>
<p>MLA tackles this:</p>
<ul>
<li>Low-Rank Compression: Instead of caching all K/V tensors at full precision, DeepSeek projects them into a smaller latent space (c_t), drastically reducing storage (sometimes &gt;90% memory savings).</li>
<li>Query Compression: Although queries are ephemeral, MLA also compresses them slightly to maintain consistency without burdening cache memory.</li>
<li>Separation of Content and Position: Rotary Positional Embeddings (RoPE) are applied independent of content representations. Keys are split into a compressed content component and a RoPE-aware positional component, ensuring positional info is preserved while keeping most KV tensors compressed.</li>
<li>Decompression-on-Demand: During attention computation, latent representations are projected back to full dimensions through a lightweight reconstruction, effectively providing the same representational power at a fraction of the memory cost.</li>
<li>Inference Impact: MLA enables long-context inference while keeping GPU memory requirements practical. In V3, KV cache memory was reduced by over 93% without degrading quality.</li>
</ul>
<h3>3. Cache Optimization Strategies</h3>
<p>DeepSeek’s architectural approach not only compresses KV tensors but complements this with efficient caching strategies:</p>
<p>Absorb vs Naive Modes:</p>
<ul>
<li>Naive: Stores all expanded K/V tensors. Highly memory-intensive.</li>
<li>Absorb: Stores only compressed latent representations, expanding them dynamically when needed. Reduces cache usage by nearly 98–99% in practice.</li>
<li>Dynamic Context Scaling: For sequences exceeding base context lengths, attention scaling factors and positional embeddings are adjusted to preserve numerical stability while maximizing effective receptive field.</li>
</ul>
<h3>4. Synergy Between MoE and MLA</h3>
<ul>
<li>Memory-Efficient Activation: Since only a subset of experts is active for every token (MoE), the memory saved by MLA complements this sparse computation, enabling model deployments with both enormous capacity and manageable resource usage.</li>
<li>Preservation of Model Quality: Unlike grouping heads arbitrarily (e.g., GQA) or quantizing caches aggressively, MLA provides low-rank compression while still allowing attention heads to access distinct representations for diverse heads.</li>
<li>Scaling to 1T Parameters (V4): With upcoming V4 models, MoE and MLA together permit ultra-high parameter counts (1T) while reducing active parameters per token (32B) and keeping memory and compute costs feasible.</li>
</ul>
<h3>5. Practical Implications</h3>
<ul>
<li>Fine-Tuning: LoRA-like adapters can be inserted post-training into attention layers to adapt MoE and MLA without retraining the full model, balancing capacity, efficiency, and accuracy.</li>
<li>High-Context Applications: For tasks like multi-document reasoning or long-session chat, MLA ensures long token histories remain manageable in GPU memory, while MoE delivers computational efficiency.</li>
<li>Inference Cost: The combination allows DeepSeek to offer API pricing that is significantly lower than dense models of comparable capability while maintaining competitive performance.</li>
</ul>
<p>These innovations make DeepSeek AI suitable for research, enterprise applications, and long-form reasoning tasks where both speed and context depth are paramount.</p>
<h3>What can you realistically expect?</h3>
<p>I wanted to build a project that would use enough tokens to realistically benchmark typical usage for a project. As I work for a warehouse and logistics company I decided to build a project that allow users to manage typical equipment found in a warehouse.</p>
<p>I called it EquipIt and you can visit the project on GitHub: (<a href="https://github.com/techno-yeti/equipit" target="_blank" rel="noopener noreferrer">https://github.com/techno-yeti/equipit</a>).</p>
<p>For EquipIt I used less than £2 of tokens, I believe this is due to the above mentioned reasons but mostly hitting the cache.</p>
<p>I wanted to show some screenshots for the EquipIt app to showcase the style and detail the AI came up with. I believe that DeepSeek v4 Pro is usable for commercial use (see project screenshots below):</p>
<p><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771497511.png" alt="Planned Requests Dashboard" /></p>
<p><strong>Planned Requests Dashboard:</strong> shows the current requests planned for that day (the day can be changed with the calendar select widget).</p>
<p><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771503792.png" alt="Listing All Requests" /></p>
<p><strong>Listing All Requests:</strong> shows all requests stored in the app (good for planning days and weeks ahead).</p>
<p><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771512375.png" alt="Printable Delivery Advice Notes" /></p>
<p><strong>Printable Delivery Advice Notes:</strong> shows printable delivery advice note for the driver to give to the security gate on exiting the yard, this can only be used once).</p>
<p><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771508039.png" alt="Limiting Email Addresses" /></p>
<p><strong>Limiting Email Addresses:</strong> to stop other companies using the app there is an email restriction function.</p>
<p>There are other features that are not shown here: trailer templating, listing equipment types, user management and the help center.</p>
<h3>Conclusion</h3>
<p>DeepSeek’s architectural elegance lies in orchestrating sparsity (MoE) and efficient attention caching (MLA) to enable high-parameter, long-context LLMs that remain economically and operationally feasible. By compressing what needs to be stored and sparsifying what needs to be computed, DeepSeek achieves a remarkable balance between scale, speed, and accuracy, representing a state-of-the-art approach to large language model design.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1785666118144.jpg" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>How to get the most out of your AI model</title>
      <link>https://julianmoors.net/posts/how-to-get-the-most-out-of-your-ai-model</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/how-to-get-the-most-out-of-your-ai-model</guid>
      <pubDate>Mon, 15 Jun 2026 18:09:38 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>How do you prevent your AI assistant wasting tokens by generating something you'll probably never use? Much like when you're sat in front of a computer deciding what to concentrate your time on, AI assistants need to be told up front what they should be focusing on. Left to their own devices, they will happily generate pages of confident-sounding code and text that simply don't fit your project.</p>
<p>After having launched my own blog and an equipment management web app as practice projects, I'll tell you how to get the most out of your AI assistant.</p>
<p>It all boils down to defining exactly what you want the AI assistant to do. By doing this, the AI is laser focused on the task at hand and you'll no longer have to worry about token burn rate. But what would this look like in practice?</p>
<p>For me, it looks like a set of spec files. Before I let an AI agent anywhere near the code, I write down exactly what the feature should do, how it should behave, and what success looks like. That little bit of upfront effort saves a surprising amount of wasted output later.</p>
<p>I have posted a set of spec files for a MERN-based project on my GitHub account (link can be found above next to my avatar picture). Feel free to edit them for your own projects, and if you would like to, post a link to an update in the comments section below.</p>
<p>Bearing in mind I used this format before I heard of agent skills. The downside is you'll have to change the spec files for each individual project. That is fine when you are working on a handful of things, but it gets repetitive if you juggle many projects.</p>
<p>I've also gotten into the habit of getting the AI model I'm working with to generate a set of spec files in case I wish to regenerate the source code with certain changes (for example if a client decides to go with an alternative CSS framework). I usually include this with the project so that I don't lose context. It means I can rebuild the whole thing from scratch weeks later without having to re-explain everything.</p>
<p>If you're looking for something more generic that can work across multiple projects, then this GitHub repo might be something to test (not created by me):</p>
<p><a href="https://github.com/addyosmani/agent-skills" target="_blank" rel="noopener noreferrer">https://github.com/addyosmani/agent-skills</a></p>
<p>The key takeaway is simple: give your AI a clear brief and it will reward you with focused, useful work. Skip that step and you will spend far more time (and tokens) cleaning up the mess.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1785666094412.jpg" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>How I used AI to roll my own blogging platform</title>
      <link>https://julianmoors.net/posts/how-i-used-ai-to-roll-my-own-blogging-platform</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/how-i-used-ai-to-roll-my-own-blogging-platform</guid>
      <pubDate>Thu, 04 Jun 2026 20:06:57 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>I took the decision to roll my own platform after wanting to experiment with agentic coding.</p>
<p>My experience of using AI has always been with using an IDE with it being used as nothing more than code suggestions on steroids. At the time of deciding to use an AI coding tool I also decided to try a new text editor.</p>
<p>After installing Zed I clicked around to get an overall feel of what Zed is capable of and I saw an agent panel indicating that Zed was giving AI agents first class support in their software.</p>
<p><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771462768.png" alt="Zed editor showing AI panel" /></p>
<p>I then signed up for Zed Pro to enable the agents selection panel and decided to use Claude Sonnet as this AI model was recommended by Zed. I saw very quickly that Claude Sonnet was eating up a chunk of money and I looked at other AI models that I thought performed just as good, but were cheaper and that's when I tried Gemini 3.5 Flash.</p>
<p>(see below the pricing table from Zed Industries' website, please note that these prices are now out dated)</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>Provider</th>
<th>Token Type</th>
<th>Provider Price per 1M tokens</th>
<th>Zed Price per 1M tokens</th>
</tr>
</thead>
<tbody><tr>
<td>Claude Opus 4.5</td>
<td>Anthropic</td>
<td>Input</td>
<td>$5.00</td>
<td>$5.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Output</td>
<td>$25.00</td>
<td>$27.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Write</td>
<td>$6.25</td>
<td>$6.875</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Read</td>
<td>$0.50</td>
<td>$0.55</td>
</tr>
<tr>
<td>Claude Opus 4.6</td>
<td>Anthropic</td>
<td>Input</td>
<td>$5.00</td>
<td>$5.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Output</td>
<td>$25.00</td>
<td>$27.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Write</td>
<td>$6.25</td>
<td>$6.875</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Read</td>
<td>$0.50</td>
<td>$0.55</td>
</tr>
<tr>
<td>Claude Opus 4.7</td>
<td>Anthropic</td>
<td>Input</td>
<td>$5.00</td>
<td>$5.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Output</td>
<td>$25.00</td>
<td>$27.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Write</td>
<td>$6.25</td>
<td>$6.875</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Read</td>
<td>$0.50</td>
<td>$0.55</td>
</tr>
<tr>
<td>Claude Opus 4.8</td>
<td>Anthropic</td>
<td>Input</td>
<td>$5.00</td>
<td>$5.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Output</td>
<td>$25.00</td>
<td>$27.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Write</td>
<td>$6.25</td>
<td>$6.875</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Read</td>
<td>$0.50</td>
<td>$0.55</td>
</tr>
<tr>
<td>Claude Sonnet 4.5</td>
<td>Anthropic</td>
<td>Input</td>
<td>$3.00</td>
<td>$3.30</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Output</td>
<td>$15.00</td>
<td>$16.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Write</td>
<td>$3.75</td>
<td>$4.125</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Read</td>
<td>$0.30</td>
<td>$0.33</td>
</tr>
<tr>
<td>Claude Sonnet 4.6</td>
<td>Anthropic</td>
<td>Input</td>
<td>$3.00</td>
<td>$3.30</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Output</td>
<td>$15.00</td>
<td>$16.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Write</td>
<td>$3.75</td>
<td>$4.125</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Read</td>
<td>$0.30</td>
<td>$0.33</td>
</tr>
<tr>
<td>Claude Haiku 4.5</td>
<td>Anthropic</td>
<td>Input</td>
<td>$1.00</td>
<td>$1.10</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Output</td>
<td>$5.00</td>
<td>$5.50</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Write</td>
<td>$1.25</td>
<td>$1.375</td>
</tr>
<tr>
<td></td>
<td>Anthropic</td>
<td>Input - Cache Read</td>
<td>$0.10</td>
<td>$0.11</td>
</tr>
<tr>
<td>GPT-5.5 pro</td>
<td>OpenAI</td>
<td>Input</td>
<td>$30.00</td>
<td>$33.00</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$180.00</td>
<td>$198.00</td>
</tr>
<tr>
<td>GPT-5.5</td>
<td>OpenAI</td>
<td>Input</td>
<td>$5.00</td>
<td>$5.50</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$30.00</td>
<td>$33.00</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Cached Input</td>
<td>$0.50</td>
<td>$0.55</td>
</tr>
<tr>
<td>GPT-5.4 pro</td>
<td>OpenAI</td>
<td>Input</td>
<td>$30.00</td>
<td>$33.00</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$180.00</td>
<td>$198.00</td>
</tr>
<tr>
<td>GPT-5.4</td>
<td>OpenAI</td>
<td>Input</td>
<td>$2.50</td>
<td>$2.75</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$15.00</td>
<td>$16.50</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Cached Input</td>
<td>$0.025</td>
<td>$0.0275</td>
</tr>
<tr>
<td>GPT-5.3-Codex</td>
<td>OpenAI</td>
<td>Input</td>
<td>$1.75</td>
<td>$1.925</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$14.00</td>
<td>$15.40</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Cached Input</td>
<td>$0.175</td>
<td>$0.1925</td>
</tr>
<tr>
<td>GPT-5.2</td>
<td>OpenAI</td>
<td>Input</td>
<td>$1.75</td>
<td>$1.925</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$14.00</td>
<td>$15.40</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Cached Input</td>
<td>$0.175</td>
<td>$0.1925</td>
</tr>
<tr>
<td>GPT-5.2-Codex</td>
<td>OpenAI</td>
<td>Input</td>
<td>$1.75</td>
<td>$1.925</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$14.00</td>
<td>$15.40</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Cached Input</td>
<td>$0.175</td>
<td>$0.1925</td>
</tr>
<tr>
<td>GPT-5 mini</td>
<td>OpenAI</td>
<td>Input</td>
<td>$0.25</td>
<td>$0.275</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$2.00</td>
<td>$2.20</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Cached Input</td>
<td>$0.025</td>
<td>$0.0275</td>
</tr>
<tr>
<td>GPT-5 nano</td>
<td>OpenAI</td>
<td>Input</td>
<td>$0.05</td>
<td>$0.055</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Output</td>
<td>$0.40</td>
<td>$0.44</td>
</tr>
<tr>
<td></td>
<td>OpenAI</td>
<td>Cached Input</td>
<td>$0.005</td>
<td>$0.0055</td>
</tr>
<tr>
<td>Gemini 3.1 Pro</td>
<td>Google</td>
<td>Input</td>
<td>$2.00</td>
<td>$2.20</td>
</tr>
<tr>
<td></td>
<td>Google</td>
<td>Output</td>
<td>$12.00</td>
<td>$13.20</td>
</tr>
<tr>
<td>Gemini 3.5 Flash</td>
<td>Google</td>
<td>Input</td>
<td>$1.50</td>
<td>$1.65</td>
</tr>
<tr>
<td></td>
<td>Google</td>
<td>Output</td>
<td>$9.00</td>
<td>$9.90</td>
</tr>
<tr>
<td>Gemini 3 Flash</td>
<td>Google</td>
<td>Input</td>
<td>$0.50</td>
<td>$0.55</td>
</tr>
<tr>
<td></td>
<td>Google</td>
<td>Output</td>
<td>$3.00</td>
<td>$3.30</td>
</tr>
</tbody></table>
<p>For those who have yet to try out AI agents, I would estimate that Gemini 3.5 Flash costs ~£50 per day for 8 hours of continuous use. Bearing in mind users won't be running AI models continuously due to code refactoring and evaluating features, but I think ~£50 per day is acceptable for fair use.</p>
<p>I hope that better frontier AI models will make the likes of Gemini 3.5 Flash cheaper in the future and perhaps even offer capped pricing at 4-6 hours of continuous use for £100 per month.</p>
<p>I guess you're wondering what features I planned for, right? Although I prompted Gemini 3.5 Flash with features, it came up with some interesting implementations:</p>
<p>The first one was a Google SEO Optimizer.<br /><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771469329.png" alt="Google SEO Optimizer" /></p>
<p>The second one was a cookies banner directly generated from its own CORS code.<br /><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771484277.png" alt="Cookies Banner" /></p>
<p>The third one was a cover image uploader, which has allowed me to re-use already existing images should I have a long running post series, for example writing a weekly news post.<br /><img src="https://blogit.lon1.cdn.digitaloceanspaces.com/media/media-6a20bd19eccfa965950a065e-1786771477925.png" alt="Cover Image Uploader" /></p>
<p>There were many other features added over time...</p>
<p>You're probably wondering if I would recommend using agentic AI for coding projects...</p>
<p>Yes, but it comes with a couple of caveats, the output code is around 95% useful, meaning the AI agent will write comments very similar to the way humans would, this makes sense if the AI models have been trained on OSS projects being hosted on places like GitHub, also you need to have a good idea of what features your app would need and this can only come from real world experience.</p>
<p>To prevent token overburn I would start off by implementing small projects at first and build up to more complicated projects after you become confident with the tools. 😉</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1785666059503.jpg" length="0" type="image/jpeg" />
    </item>
    <item>
      <title>Hello World</title>
      <link>https://julianmoors.net/posts/hello-world</link>
      <guid isPermaLink="true">https://julianmoors.net/posts/hello-world</guid>
      <pubDate>Wed, 03 Jun 2026 23:59:26 GMT</pubDate>
      <author>julian.moors@outlook.com (Julian Moors)</author>
      <description><![CDATA[<p>Welcome to my blog.</p>
<p>Here I write mostly about my experiences with coding and AI tools.</p>
<p>This blogging software was created solely with agentic AI models. In a later post I describe how I used AI to design and implement this project and the reasons why I chose this method.</p>
<p>If you have yet to dip your toes in Agentic Driven Development then I hope to inspire you to start your own projects.</p>
]]></description>
      <enclosure url="https://julianmoors.nethttps://blogit.lon1.cdn.digitaloceanspaces.com/posts/post-6a20bd19eccfa965950a065e-1785666034994.jpg" length="0" type="image/jpeg" />
    </item>
  </channel>
</rss>