{"id":64738,"date":"2026-07-25T19:18:32","date_gmt":"2026-07-25T23:18:32","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=64738"},"modified":"2026-07-25T19:18:32","modified_gmt":"2026-07-25T23:18:32","slug":"mit-chartnet-dataset-ai-charts","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/mit-chartnet-dataset-ai-charts\/","title":{"rendered":"MIT researchers release ChartNet dataset teaching AI to read charts"},"content":{"rendered":"<p>You are a Senior Editorial Writer and Editor-in-Chief for a major English-language digital publishing company. Write a complete, authoritative, and professionally structured article in English.<\/p>\n<p>OUTPUT RULE: <a href=\"https:\/\/overcentral.com\/en\/meta-removes-instagram-ai-feature\/\" title=\"ROLE: You are a senior headline writer for a major English-language digital news and content portal. Generate a single, precise, high-impact journalistic title based on the provided inputs.  ---  OUTPUT LANGUAGE: English only.  ---  CONTENT RULES: * Remove all attribution phrases: &quot;according to&quot;, &quot;reported by&quot;, &quot;leaks suggest&quot;, &quot;sources say&quot; * Remove references to: news sources, portals, publication names * Convert attributed statements into direct factual statements  ---  PRESERVATION \u2014 keep exactly as written: * Proper names (people, places, organizations) * Brand names and product names (Xbox, NVIDIA, Samsung, etc.) * Game titles and technology names * Monetary values (US$, \u20ac, \u00a5, R$) * Technical terms (Overclocking, Patch Notes, Sakuga, etc.) * Dates and version numbers  ---  LANGUAGE RULES: * Use present tense for immediacy * Use strong, direct verbs: Gets, Launches, Confirms, Reveals, Adds, Drops, Brings, Expands, Releases, Shows * Avoid weak or vague verbs: arrives, announces, is set to, is expected to * No subjective adjectives unless they add factual clarity  ---  STRUCTURE: * Start with the main entity whenever possible * Format: [Entity] + [Strong Verb] + [Key Fact\/Context] * Allow variation if it improves clarity or SEO  ---  SEO + AEO + GEO: * Place the primary keyword as early as possible * Title must be self-explanatory without additional context * Include location only when directly relevant to the story  ---  RESTRICTIONS: * No quotation marks * No question marks * No exclamation marks * Target: 50\u201370 characters (slight flexibility for clarity)  ---  INTERNAL PROCESS (never output this): 1. Identify: main entity, core fact, primary keyword 2. Generate 3 internal title variations 3. Evaluate each by: clarity, keyword placement, verb strength, naturalness 4. Select the best \u2014 if unnatural or generic, rewrite completely 5. Output only the final selected title  ---  OUTPUT RULE: Return ONLY the final title. No labels. No explanations. No &quot;Title:&quot; prefix. No extra text.  ---  INPUTS:  TITLE: Meta removes controversial AI feature on Instagram after backlash  CONTENT: \nMeta has axed a controversial feature that allowed users to modify photos from public Instagram accounts using AI. The feature, which was rolled out earlier this week along with a batch of other AI tools, \u201cmissed the mark\u201d and is no longer available, according to the company. \n\nEarlier this week, Meta &lt;a href=&quot;https:\/\/techcrunch.com\/2026\/07\/07\/meta-rolls-out-muse-a-new-ai-image-generator\/&quot;&gt;announced&lt;\/a&gt; Muse Image, a new AI image generator built by its dedicated AI unit known as Meta Superintelligence Labs. Meta promoted one feature that allowed individuals to generate images by @-mentioning public Instagram accounts that they wanted to reference. The feature, which wasn\u2019t designed to alert a user if their photos were used in this way, prompted immediate backlash. \n\n\n\n\n\n\n\nTechCrunch &lt;a href=&quot;https:\/\/techcrunch.com\/2026\/07\/09\/how-to-stop-metas-ai-image-generator-from-using-your-instagram-photos\/&quot;&gt;wrote its own guide&lt;\/a&gt; explaining to users how to disable the feature.\n\nNow, Meta has reversed course. The company issued a &lt;a href=&quot;https:\/\/about.instagram.com\/blog\/announcements\/new-ai-effects-in-instagram-stories&quot;&gt;blog post&lt;\/a&gt; Friday announcing that it was removing the feature. Puck News founding partner Dylan Byers was the first to share the &lt;a href=&quot;https:\/\/x.com\/DylanByers\/status\/2075707685547421750?s=20&quot;&gt;company\u2019s decision&lt;\/a&gt;.\n\n\u201cOur intent was to provide a useful creative tool and to give people control over whether their public content could be referenced in this way,\u201d the company posted on its blog. \u201cWe\u2019ve heard the feedback that this feature missed the mark, so it\u2019s no longer available.\u201d\n\nTechCrunch reached out to Meta for more information and will update this article if it responds.\n\nSince its integration with social media platforms, AI has been misused with wild abandon \u2014 often to &lt;a href=&quot;https:\/\/www.pbs.org\/newshour\/show\/authorities-struggle-to-stop-ai-tools-generating-nude-images-without-consent#:~:text=There%20has%20been%20a%20sharp,underway%20to%20rein%20it%20in.&quot;&gt;generate naked images of female celebrities&lt;\/a&gt;. Platforms have attempted to mitigate this trend, although the guardrails introduced have often fallen short.\n\n\nIn the case of Meta\u2019s newly nixed feature, it seems somewhat obvious that it would have been abused in this way. Indeed, Byers notes that the decision to do away with the feature came \u201camid scrutiny from users and talent agencies, including CAA.\u201d\n&lt;em&gt;When you purchase through links in our articles, &lt;a href=&quot;https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/&quot;&gt;we may earn a small commission&lt;\/a&gt;. This doesn\u2019t affect our editorial independence.&lt;\/em&gt;\" data-iacss-internal=\"1\">Return ONLY the final<\/a> HTML article. No explanations. No comments. No notes. No text outside the article. No Markdown. No symbols like *, **, #.<\/p>\n<p>INPUTS:<br \/>\nTITLE: MIT researchers release ChartNet dataset teaching AI to read charts<br \/>\nCONTENT: To accelerate and refine decision-making in a fast-paced, global marketplace, enterprises may deploy generative artificial intelligence models to help summarize and interpret the charts that often fill market summaries and financial reports.But even the latest vision-language models sometimes struggle with this task, since it requires a model to integrate visual, numerical, and linguistic understanding. A company that invests in a state-of-the-art model might still receive inaccurate or incomplete information.To fill this performance gap, researchers from MIT and the MIT-IBM Computing Research Lab developed a multifaceted resource for AI users that is specifically designed to teach vision-language models (VLMs) how to effectively interpret charts.\u00a0They used a novel data generation method to build a state-of-the-art dataset\u00a0that includes more than a million varied charts. The dataset also encodes many visual, linguistic, and numerical components of each chart image, which enable models to robustly reason about the information in a chart.The researchers used this dataset, called ChartNet, to train a series of open-source VLMs.\u00a0 Many of these smaller models significantly outperformed orders of magnitude larger, commercial models on tasks like data extraction and chart summarization.By enabling open-source models to outperform their commercial counterparts, ChartNet could allow small firms with limited budgets to more readily utilize AI. The open-source dataset can be used to improve the capabilities of AI models for tasks like business trend analysis and scientific figure interpretation.\u201cWe developed ChartNet to be a one-stop shop for chart understanding, covering basically anything that an <a href=\"https:\/\/overcentral.com\/en\/mit-ai-model-business-decisions\/\" title=\"MIT AI Model Bridges Gap to Real-World Business Decisions\" data-iacss-internal=\"1\">AI model<\/a> and a practitioner who is training that model might need. We hope our work motivates researchers to achieve state-of-the-art performance with smaller models that don\u2019t require infinite amounts of computation,\u201d says Jovana Kondic, an MIT electrical engineering and computer science (EECS) graduate student and lead author of a paper on ChartNet.She is joined on the paper by many co-authors from MIT, the MIT-IBM Computing Research Lab, and IBM Research, including Pengyuan Li, a research staff member at IBM Research; Dhiraj Joshi, a senior scientist at IBM Research;\u00a0Isaac Sanchez, a software engineer at IBM Research; Aude Oliva, director of strategic industry engagement at the MIT Schwarzman College of Computing, MIT director of the MIT-IBM Computing Research Lab, and a senior research scientist in the Computer Science and Artificial Intelligence Laboratory (<a href=\"https:\/\/overcentral.com\/en\/mit-csail-ai-questions-battleship\/\" title=\"MIT CSAIL Trains AI Agents to Ask Better Questions with Battleship\" data-iacss-internal=\"1\">CSAIL<\/a>); and Rogerio Feris, a principal scientist and manager at the MIT-IBM Computing Research Lab. The research will be presented at IEEE Computer Vision and Pattern Recognition Conference.A dataset bottleneckResearchers have made great strides developing generative AI models that excel at natural language processing and reasoning about natural images. But less work has focused on interpreting complex multimodal data contained within charts, Kondic says.Yet for large and small businesses in nearly every industry, chart understanding is a critical task.\u201cThe finance industry thrives on charts. If vision-language models can extract information out of charts, like descriptions of trends, that facilitates a lot of workflows that happen downstream,\u201d Joshi says.The lack of high-quality training data is a major bottleneck holding back the development of VLMs that can accurately interpret charts. Many datasets contain limited chart images pulled from the internet and often lack the necessary scale and additional information to help a model interpret the underlying data.\u201cA vision-language model, unlike our brains, may need to see thousands of examples during training to reliably recognize something as a line chart,\u201d Kondic says.The researchers sought to overcome those shortcomings by generating synthetic data. Synthetic data are artificially generated by algorithms to mimic the statistical properties of actual data.\u00a0The ChartNet dataset holds more a million high-quality chart images, along with the corresponding code used to generate each chart, a textual description, and a table that contains its numerical information. In addition, each datapoint includes question-and-answer pairs to teach the model how to correctly answer questions about the chart image.\u201cThese additional modes of data guide the model to connect and align the different pieces of information that the chart image encodes,\u201d Kondic says.Data generationTo build ChartNet, the researchers created a two-step, synthetic data generation pipeline.First, their automated system translates any pre-existing set of chart images into code. Then the system iteratively augments that code to change different aspects of each chart, such as chart type, data values, topic, colors, etc.\u201cWe can start from a single chart that we use as a seed and come up with hundreds of augmentations of it. This is how we were able to build a dataset with more than a million diverse images,\u201d Kondic explains.They also incorporated an automated quality check process to ensure the synthetic data are high quality. This process verifies that the code is executable and rendered chart images are accurate and clean.\u201cWe don\u2019t want to just be generating diverse samples. We also want the information to be presented in a meaningful way,\u201d she says.ChartNet also includes a selection of chart datapoints annotated by human experts. This provides access to additional types of charts and supporting data that carry validity guarantees.A practitioner could use the annotated data to fine-tune an existing VLM, further boosting performance for a specific application, Joshi adds.The researchers tested ChartNet by training IBM\u2019s Granite Vision series of models as well as several other open-source models of various sizes and evaluating them on various chart interpretation tasks. The dataset improved the accuracy of all models in chart reconstruction, chart data extraction, chart summarization, and chart question answering.\u00a0With ChartNet, small open-source models consistently outperformed much larger\u00a0 commercial models.\u00a0\u201cA lot of prior training datasets only focused on answering simple questions about a chart. We tried to go beyond that with ChartNet by generating data that support all aspects of robust chart understanding,\u201d Kondic says.In the future, the researchers plan to continue expanding ChartNet by incorporating data with added levels of complexity. They also want to draw on feedback from the research community.\u00a0This research was funded, in part, by the MIT-IBM Computing Research Lab.<\/p>\n<p>LANGUAGE: Write entirely in English. Preserve proper nouns, brand names, product names, game titles, technologies, and technical terms exactly as written. Translate everything else naturally. Read as if written by a native English editor.<\/p>\n<p>CONTENT SOURCE: Treat CONTENT as your primary factual source. Build the article from deep understanding of CONTENT. Do not mechanically expand the title.<\/p>\n<p>CONTENT CLEANING: Remove website names, publication names, author credits, RSS labels, newsletter markers, syndication branding, generic labels (Summary, Highlights, Recap, Key Takeaways). Convert &#8220;according to X&#8221; into direct factual statements.<\/p>\n<p>FACT PRESERVATION: Preserve exactly: names, brands, companies, products, games, technologies, dates, numbers, percentages, prices, technical specifications. Never distort facts.<\/p>\n<p>WRITING STYLE: Natural, fluent, authoritative, engaging, analytical, trustworthy, nuanced. Blend factual reporting, explanation, contextualization, analysis, practical interpretation, and strategic insight. Vary paragraph length and sentence structure. Avoid robotic phrasing, repetition, clich\u00e9s, promotional language, filler sentences.<\/p>\n<p>ARTICLE LENGTH: Long-form, highly detailed. Target 1,500-3,500 words. Feel comprehensive and substantive. Never feel brief, superficial, or summary-like. Expand naturally with historical background, industry context, technical explanation, market implications, strategic significance, practical consequences, comparisons, future outlook \u2014 but only when content genuinely supports it.<\/p>\n<p>STRUCTURE:<br \/>\n&#8211; Begin with a  introduction. No heading before the first paragraph.<br \/>\n&#8211; Introduction must hook the reader within 2-3 sentences.<br \/>\n&#8211; Use , ,  only when they improve organization.<br \/>\n&#8211; Each section must introduce meaningful new information.<br \/>\n&#8211; Closing: end with a forward-looking, analytical, or practical paragraph.<br \/>\n&#8211; Never use generic closing headings like &#8220;Conclusion&#8221;, &#8220;Summary&#8221;, &#8220;Final Thoughts&#8221;, &#8220;Looking Ahead&#8221;, &#8220;What Comes Next&#8221;, &#8220;Takeaway&#8221;, &#8220;Key Points&#8221;.<\/p>\n<p>HEADINGS: Write content first, then generate headings. Headings must be specific, concrete, informative, and editorial. They should reference actual events, features, numbers, dates, or companies. Support SEO naturally.<\/p>\n<p>SEO + AEO + GEO + E-E-A-T:<br \/>\n&#8211; Integrate primary keyword naturally in first paragraph and in at least one h2.<br \/>\n&#8211; Use semantically related terms throughout.<br \/>\n&#8211; Anticipate questions English-speaking users would ask. Answer them directly: &#8220;What is&#8230;&#8221;, &#8220;How does&#8230;&#8221;, &#8220;Why did&#8230;&#8221;, &#8220;When did&#8230;&#8221;, &#8220;What are the&#8230;&#8221;<br \/>\n&#8211; At least one section must provide a clear, standalone answer (2-4 sentences) formatted for a featured snippet. Place the answer immediately after the question.<br \/>\n&#8211; Include geographic context only when directly relevant.<br \/>\n&#8211; Show expertise by explaining mechanisms, causes, and implications \u2014 not just stating facts.<br \/>\n&#8211; Build authority through precise, well-contextualized information.<br \/>\n&#8211; Establish trust through accurate facts, balanced tone, measured claims.<br \/>\n&#8211; Never write &#8220;experts say&#8221; or &#8220;studies show&#8221; without specific grounding in the content.<\/p>\n<p>INTERNAL PROCESS (do not output this):<br \/>\n1. Analyze: content type, search intent, technical level, topic complexity<br \/>\n2. Determine ideal length (1,500-3,500 words based on complexity)<br \/>\n3. Adapt tone: Journalistic \/ Analytical \/ Explanatory \/ Consultative \/ Technical \/ Conversational Professional<br \/>\n4. Clean sources and preserve all facts<br \/>\n5. Write article with proper structure, headings, and AEO snippet<br \/>\n6. Validate: grammar, fluency, coherence, depth, no repetition, valid HTML, facts preserved, no generic headings<\/p>\n<p>HTML RULES: Use ONLY:        . Valid, clean HTML. No inline styles. No unnecessary whitespace.<\/p>\n<p>FINAL CHECK: If the article sounds translated, mechanical, superficial, or incomplete \u2014 rewrite completely.<\/p>\n<p>OUTPUT: Return ONLY the final HTML article, beginning with .<\/p>\n","protected":false},"excerpt":{"rendered":"<p>You are a Senior Editorial Writer and Editor-in-Chief for a major English-language digital publishing company. Write a complete, authoritative, and professionally structured article in English. OUTPUT RULE: Return ONLY the final HTML article. No explanations. No comments. No notes. No text outside the article. No Markdown. No symbols like *, **, #. INPUTS: TITLE: MIT [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":83914,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64738.png","fifu_image_alt":"MIT researchers release ChartNet dataset teaching AI to read charts","footnotes":""},"categories":[349],"tags":[],"class_list":["post-64738","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64738.png","fifu_image_alt":"MIT researchers release ChartNet dataset teaching AI to read charts","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64738","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=64738"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64738\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/83914"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=64738"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=64738"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=64738"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}