{"id":8771,"date":"2025-07-20T07:30:53","date_gmt":"2025-07-20T04:30:53","guid":{"rendered":"https:\/\/handoli.com\/index.php\/2025\/07\/20\/data-machina-257\/"},"modified":"2025-07-20T07:30:53","modified_gmt":"2025-07-20T04:30:53","slug":"data-machina-257","status":"publish","type":"post","link":"https:\/\/handoli.com\/index.php\/2025\/07\/20\/data-machina-257\/","title":{"rendered":"Data Machina #257"},"content":{"rendered":"<p><strong>On Compound AI Systems, Txt2SQL &amp; Data Agents.  <\/strong>A year ago or so, a client enthusiastically presented us with a long list of \u201c<em>AI LLM projects<\/em>.\u201d Among them, there was one project listed as: \u201c<em>use text-to-sql to automate all data analysis tasks<\/em>\u201d  \u2026 We thought: \u201cUmm\u2026 this is going to be an <em>interesting<\/em> project\u201d\u2026 Months later the client abandoned the project. <\/p>\n<p><strong>Pioneers in Text-to-SQL at enterprise scale<\/strong>. afaik, Pinterest was one of the first companies that deployed Tex2SQL at scale in enterprise production. Importantly, they were one of the first ones in sharing their experience. This is an excellent post in which the engineering team describes the whole journey from academic Txt2SQL to production. Blogpost: <a href=\"https:\/\/medium.com\/pinterest-engineering\/how-we-built-text-to-sql-at-pinterest-30bad30dabff\">How we built Text-to-SQL at Pinterest<\/a>.<\/p>\n<div class=\"captioned-image-container\">\n<figure><a class=\"image-link image2 is-viewable-img\" target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/%24s_!e-G-!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc597df-c52e-4ffd-89ec-5767cfdd6f54_2616x1794.png\" data-component-name=\"Image2ToDOM\">\n<div class=\"image2-inset\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/%24s_!e-G-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc597df-c52e-4ffd-89ec-5767cfdd6f54_2616x1794.png\" width=\"452\" height=\"309.81868131868134\" data-attrs='{\"src\":\"https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/6fc597df-c52e-4ffd-89ec-5767cfdd6f54_2616x1794.png\",\"srcNoWatermark\":null,\"fullscreen\":null,\"imageSize\":null,\"height\":998,\"width\":1456,\"resizeWidth\":452,\"bytes\":795366,\"alt\":null,\"title\":null,\"type\":\"image\/png\",\"href\":null,\"belowTheFold\":false,\"topImage\":true,\"internalRedirect\":null,\"isProcessing\":false,\"align\":null,\"offset\":false}' class=\"sizing-normal\" alt=\"\"\/>\n<div class=\"image-link-expand\">\n<div class=\"pencraft pc-display-flex pc-gap-8 pc-reset\">\n<div class=\"pencraft pc-reset icon-container restack-image\"><\/div>\n<div class=\"pencraft pc-reset icon-container view-image\"><\/div>\n<\/div>\n<\/div>\n<\/div>\n<p><\/p><\/a><\/figure>\n<\/div>\n<p><strong>Text-to-SQL augmented with RAG: Not easy yet<\/strong>. Sure! LLMs can write beautiful, syntactically correct SQL statements because there are tons of public SQL code on which LLMs have been trained on. But LLMs have quite a bit of challenges when dealing with real-world relational data problems. <a href=\"https:\/\/blog.getwren.ai\/4-key-technical-challenges-using-rag-with-llms-to-query-database-text-to-sql-and-how-to-solve-it-5d5a3d6682e5\">Top 4 Challenges using RAG with LLMs to Text-to-SQL and how to solve it<\/a>.<\/p>\n<p><strong>Text-to-SQL or human-like AI analysts?<\/strong> This is an interesting post by the team at Pattern, a startup building a financial analysis agent. Their main idea: Text-to-SQL should be more like Text-to-Analysis that works at the business layer. And the LLM -beyond prompting- should behave like a human analyst by using multiple, specialist AI agents that contribute to the analysis process. Blogpost: <a href=\"https:\/\/patterns.app\/blog\/text-to-sql-and-its-uncanny-valley\">Text to SQL and its uncanny valley.<\/a> <\/p>\n<p><strong>Agents for data workflows<\/strong>. In the real world, data workflows with several data pipelines with messy, dirty, changing data are an absolute nightmare! Enter Meadow: An open source, agentic framework for building multi-agent data workflows with LLMs with interactive user feedback. Meadow\u2019s approach is to chain several specialised agents like Text-to-SQL, Planner, Executor, Schema Cleaner, Validator, Router agents to perform an end to end data workflow. Blogpost and repo here: <a href=\"https:\/\/numbersstation.ai\/introducing-meadow-llm-agents-for-data-tasks\/\">Introducing Meadow: LLM Agents for Data Tasks<\/a>. <\/p>\n<div class=\"captioned-image-container\">\n<figure><a class=\"image-link image2 is-viewable-img\" target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/%24s_!7yW0!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d615d67-74d5-497c-a555-52e0f61d18b2_1514x886.png\" data-component-name=\"Image2ToDOM\">\n<div class=\"image2-inset\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/%24s_!7yW0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d615d67-74d5-497c-a555-52e0f61d18b2_1514x886.png\" width=\"540\" height=\"315.989010989011\" data-attrs='{\"src\":\"https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/4d615d67-74d5-497c-a555-52e0f61d18b2_1514x886.png\",\"srcNoWatermark\":null,\"fullscreen\":null,\"imageSize\":null,\"height\":852,\"width\":1456,\"resizeWidth\":540,\"bytes\":204938,\"alt\":null,\"title\":null,\"type\":\"image\/png\",\"href\":null,\"belowTheFold\":false,\"topImage\":false,\"internalRedirect\":null,\"isProcessing\":false,\"align\":null,\"offset\":false}' class=\"sizing-normal\" alt=\"\"\/>\n<div class=\"image-link-expand\">\n<div class=\"pencraft pc-display-flex pc-gap-8 pc-reset\">\n<div class=\"pencraft pc-reset icon-container restack-image\"><\/div>\n<div class=\"pencraft pc-reset icon-container view-image\"><\/div>\n<\/div>\n<\/div>\n<\/div>\n<p><\/p><\/a><\/figure>\n<\/div>\n<p><strong>Compound AI Systems and LLM Data Agents<\/strong>. In February, researchers at Berkeley AIR published an article <a href=\"https:\/\/bair.berkeley.edu\/blog\/2024\/02\/18\/compound-ai-systems\/\">on the rapid Shift from Models to Compound AI Systems.<\/a> In contrast to a \u201cclassic\u201d rather static AI model, a Compound AI System interacts with multiple components\u2026 function calls, APIs, search, retrievers, agents\u2026 The researchers argue that you should design and implement an AI system from the perspective of a Compound AI System. <\/p>\n<p>More recently, Howard at Wren.ai &#8211; a  startup offering RAG-Txt2SQL solutions- wrote a great post <a href=\"https:\/\/blog.getwren.ai\/the-new-wave-of-composable-data-systems-and-the-interface-to-llm-agents-ec8f0a2e7141\">on the new wave and the concept of Composable Data Systems and the Interface to LLM agents<\/a>. And Mosaic AI (aka Databricks AI) just announced a series of new capabilities <a href=\"https:\/\/www.databricks.com\/blog\/mosaic-ai-build-and-deploy-production-quality-compound-ai-systems\">on building and deploying production-quality Compound AI Systems<\/a>.<\/p>\n<div class=\"captioned-image-container\">\n<figure><a class=\"image-link image2 is-viewable-img\" target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/%24s_!Kecg!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b57b7bc-9ee1-4962-86e3-1cb176ba6b37_1964x1238.png\" data-component-name=\"Image2ToDOM\">\n<div class=\"image2-inset\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/%24s_!Kecg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b57b7bc-9ee1-4962-86e3-1cb176ba6b37_1964x1238.png\" width=\"472\" height=\"297.5934065934066\" data-attrs='{\"src\":\"https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/5b57b7bc-9ee1-4962-86e3-1cb176ba6b37_1964x1238.png\",\"srcNoWatermark\":null,\"fullscreen\":null,\"imageSize\":null,\"height\":918,\"width\":1456,\"resizeWidth\":472,\"bytes\":497875,\"alt\":null,\"title\":null,\"type\":\"image\/png\",\"href\":null,\"belowTheFold\":false,\"topImage\":false,\"internalRedirect\":null,\"isProcessing\":false,\"align\":null,\"offset\":false}' class=\"sizing-normal\" alt=\"\"\/>\n<div class=\"image-link-expand\">\n<div class=\"pencraft pc-display-flex pc-gap-8 pc-reset\">\n<div class=\"pencraft pc-reset icon-container restack-image\"><\/div>\n<div class=\"pencraft pc-reset icon-container view-image\"><\/div>\n<\/div>\n<\/div>\n<\/div>\n<p><\/p><\/a><\/figure>\n<\/div>\n<p><strong>A fully open source NL engine for SQL DBs<\/strong>. Last month, the team at Dataherald open sourced natural language-to-SQL engine built for enterprise-level question answering over relational data. It allows you to set up an API from your database that can answer questions in plain English. You can do NL Q&amp;A on Prod DBs,  NL queries the DW directly without IT support, or create a ChatGPT plugin. Repo and docs here: <a href=\"https:\/\/github.com\/Dataherald\/dataherald\">Interact with your SQL database, Natural Language to SQL using LLMs<\/a>.<\/p>\n<p><strong>Open Source AI Agents for Data Analysis<\/strong>. PandasAI just open sourced a Python library that makes it easy to ask questions to your data in natural language. Beyond querying, PandasAI offers functionalities to visualize data through graphs, cleanse datasets by addressing missing values, and enhance data quality through feature generation, making it a comprehensive tool for data scientists and analysts. Checkout the repo and docs here: <a href=\"https:\/\/github.com\/Sinaptik-AI\/pandas-ai\">Pandas AI &#8211; AI agents for Data Analysis<\/a>.<\/p>\n<p><strong>Free course: Building Your Own Database Agent<\/strong>. In this course, you will develop an AI agent that interacts with databases using natural language, simplifying the process for querying and extracting insights. <a href=\"https:\/\/www.deeplearning.ai\/short-courses\/building-your-own-database-agent\/\">To know more and join this free course click here<\/a>.<\/p>\n\n<p>Have a nice week.<\/p>\n<p class=\"button-wrapper\" data-attrs='{\"url\":\"https:\/\/datamachina.substack.com\/subscribe?\",\"text\":\"Subscribe now\",\"action\":null,\"class\":null}' data-component-name=\"ButtonCreateButton\"><a class=\"button primary\" href=\"https:\/\/datamachina.substack.com\/subscribe?\"><span>Subscribe now<\/span><\/a><\/p>\n<h3><strong>10 Link-o-Troned<\/strong><\/h3>\n<ol>\n<li>\n<p><a href=\"https:\/\/readmedium.com\/en\/my-thoughts-on-apple-intelligence-16a793359cb5\">My Thoughts on Apple Intelligence Models<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/yellow-apartment-148.notion.site\/AI-Search-The-Bitter-er-Lesson-44c11acd27294f4495c3de778cd09c8d\">AI Search: The Bitter-er Lesson<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/medium.com\/@cpdough\/building-ai-agents-lessons-learned-over-the-past-year-41dc4725d8e5\">Building AI Agents: Lessons Learned over the Past Year<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/github.com\/pratyushmaini\/llm_dataset_inference\/\">[gotcha] LLM Dataset Inference: Did you Train on My Dataset? <\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/trigaten.github.io\/Prompt_Survey_Site\/\">A Systematic Survey on [the latest] Prompting Techniques, 6\/2024<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/readmedium.com\/en\/the-challenges-of-retrieving-and-evaluating-relevant-context-for-rag-e362f6eaed34\">The Challenges of Retrieving &amp; Evaluating Relevant Context for RAG<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/blogs.nvidia.com\/blog\/nemotron-4-synthetic-data-generation-llm-training\/\">NVIDIA Nemotron-4 340B: A SOTA Open Synthetic DataGen Pipeline<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/www.lamini.ai\/blog\/lamini-memory-tuning\">[new] Memory Tuning: 95% LLM Accuracy, 10x Fewer Hallucinations<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/microsoft.github.io\/generative-ai-for-beginners\/#\/\">[free, very comprehensive] Generative AI for Beginners v2, 18 Lessons<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/www.youtube.com\/watch?v=orDKvo8h71o&amp;list=PLoROMvodv4rNiJRchCzutFw5ItR_Z27CM&amp;index=33\">[great] Stanford CS25 v4: A Highly Opinionated View on Transformers<\/a><\/p>\n<\/li>\n<\/ol>\n<div>\n<hr \/>\n<\/div>\n<p class=\"button-wrapper\" data-attrs='{\"url\":\"https:\/\/datamachina.substack.com\",\"text\":\"Share Data Machina with your friends\",\"action\":null,\"class\":\"button-wrapper\"}' data-component-name=\"ButtonCreateButton\"><a class=\"button primary button-wrapper\" href=\"https:\/\/datamachina.substack.com\/\"><span>Share Data Machina with your friends<\/span><\/a><\/p>\n<div>\n<hr \/>\n<\/div>\n<h3><strong>the ML Pythonista<\/strong><\/h3>\n<ol>\n<li>\n<p><a href=\"https:\/\/www.youtube.com\/watch?v=l8pRSuU81PU\">[mega session] Karpathy\u2019s <\/a><em><a href=\"https:\/\/www.youtube.com\/watch?v=l8pRSuU81PU\">Let\u2019s Reproduce GPT-2<\/a><\/em><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/www.youtube.com\/watch?v=p0I-hwZSWMs\">Function Calling with OpenAI APIs: A Crash Course<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/github.com\/alipay\/agentUniverse\">agentUniverse &#8211; An Apache 2.0 Framework for Multi-Agent Apps<\/a><\/p>\n<\/li>\n<\/ol>\n<h3><strong>Deep &amp; Other Learning Bits<\/strong><\/h3>\n<ol>\n<li>\n<p><a href=\"https:\/\/arxiv.org\/pdf\/2406.08929\">[free] Step-by-Step Diffusion: An Elementary Tutorial (pdf)<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/machinelearning.apple.com\/research\/introducing-apple-foundation-models\">[official] Intro to Apple\u2019s On-Device &amp; Server Foundation Models<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/sakana.ai\/llm-squared\/\">Can LLMs Invent Better Ways to Train LLMs? An Auto-Evolutionary Approach<\/a><\/p>\n<\/li>\n<\/ol>\n<h3><strong>AI\/ DL ResearchDocs<\/strong><\/h3>\n<ol>\n<li>\n<p><a href=\"https:\/\/arxiv.org\/abs\/2406.09308\">DeepMind &#8211; Transformers Meet Neural Algorithmic Reasoners<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/arxiv.org\/abs\/2406.04692\">Together.ai &#8211; Mixture of Agents (MoAs) Enhances LLMs Capabilities<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/github.com\/microsoft\/Samba\">MSR &#8211; Samba: Mamba SSM + MLP + Sliding Window for Unlimited Context<\/a><\/p>\n<\/li>\n<\/ol>\n<h3><strong>MLOps Untangled<\/strong><\/h3>\n<ol>\n<li>\n<p><a href=\"https:\/\/langfuse.com\/\">LangFuse &#8211; An Open Source LLM Engineering Platform<\/a> <\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/medium.com\/nebius\/slurm-vs-kubernetes-which-to-choose-for-your-ml-workloads-23e398ce7ece\">Choosing Slurm vs Kubernetes for Modern ML Workloads<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/medium.com\/@mlengineering\/llm-monitoring-and-observability-tools-tips-and-best-practices-98ea16f533a7\">LLM Monitoring and Observability: Tools, Tips &amp; Best Practices<\/a><\/p>\n<\/li>\n<\/ol>\n<h3><strong>ML Datasets &amp; Stuff<\/strong><\/h3>\n<ol>\n<li>\n<p><a href=\"https:\/\/www.haqtu.me\/Recap-Datacomp-1B\/\">What If We Recaption Billions of Web Images with LLaMA-3?<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/arxiv.org\/abs\/2406.08673\">NVIDIA HelpSteer2: An Open Dataset for Training Top Reward Models<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/huggingface.co\/datasets\/NousResearch\/CharacterCodex\">Character Codex: A Dataset of Characters in Comics, Movies, TV Shows for GenAI<\/a><\/p>\n<\/li>\n<\/ol>\n<h3><strong>Postscript, etc <\/strong><\/h3>\n<div class=\"captioned-button-wrap\" data-attrs='{\"url\":\"https:\/\/datamachina.substack.com\/p\/data-machina-257?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share\",\"text\":\"Share\"}' data-component-name=\"CaptionedButtonToDOM\">\n<div class=\"preamble\">\n<p class=\"cta-caption\">Enjoyed this post? Tell your friends about Data Machina. Thanks for reading.<\/p>\n<\/div>\n<p class=\"button-wrapper\" data-attrs='{\"url\":\"https:\/\/datamachina.substack.com\/p\/data-machina-257?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share\",\"text\":\"Share\"}' data-component-name=\"ButtonCreateButton\"><a class=\"button primary\" href=\"https:\/\/datamachina.substack.com\/p\/data-machina-257?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share\"><span>Share<\/span><\/a><\/p>\n<\/div>\n<p>Tips? Suggestions? Feedback?\u00a0<a href=\"mailto:carlos@datamachina.com\">email Carlos<\/a><\/p>\n<p>Curated by\u00a0<a href=\"https:\/\/twitter.com\/ds_ldn\">@ds_ldn\u00a0<\/a>in the middle of the night.<\/p>","protected":false},"excerpt":{"rendered":"<p>On Compound AI Systems, Txt2SQL &amp; Data Agents. A year ago or so, a client enthusiastically presented us with a long list of \u201cAI LLM projects.\u201d Among them, there was one project listed as: \u201cuse text-to-sql to automate all data analysis tasks\u201d \u2026 We thought: \u201cUmm\u2026 this is going to be an interesting project\u201d\u2026 Months [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":8772,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rop_custom_images_group":[],"rop_custom_messages_group":[],"rop_publish_now":"initial","rop_publish_now_accounts":[],"rop_publish_now_history":[],"rop_publish_now_status":"pending","footnotes":""},"categories":[54],"tags":[],"class_list":["post-8771","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-airepos"],"_links":{"self":[{"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/posts\/8771","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/comments?post=8771"}],"version-history":[{"count":0,"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/posts\/8771\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/media\/8772"}],"wp:attachment":[{"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/media?parent=8771"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/categories?post=8771"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/handoli.com\/index.php\/wp-json\/wp\/v2\/tags?post=8771"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}