{"id":21381,"date":"2022-12-22T11:40:52","date_gmt":"2022-12-22T09:40:52","guid":{"rendered":"https:\/\/www.wjst.de\/blog\/?p=21381"},"modified":"2023-12-06T10:57:47","modified_gmt":"2023-12-06T08:57:47","slug":"from-big-data-to-good-data","status":"publish","type":"post","link":"https:\/\/www.wjst.de\/blog\/sciencesurf\/2022\/12\/from-big-data-to-good-data\/","title":{"rendered":"From big data to good data"},"content":{"rendered":"<p>I think the AI Community is slowly <a href=\"https:\/\/spectrum.ieee.org\/andrew-ng-data-centric-ai\">getting at the point<\/a> where epidemiologists had already been two decades ago, so Ng:<\/p>\n<blockquote><p>\u201cIn many industries where giant data sets simply don\u2019t exist, I think the focus has to shift from big data to good data. Having 50 thoughtfully engineered examples can be sufficient to explain to the neural network what you want it to learn.\u201d<br \/>\n\u2014Andrew Ng, CEO Landing AI<\/p><\/blockquote>\n<p>or Bickson: <a href=\"https:\/\/gradientflow.com\/large-image-datasets-today-are-a-mess\/\">Large Image Datasets Today Are a Mess<\/a><\/p>\n<blockquote><p>&#8220;We were surprised to find that there are 1.2M pairs of identical images in ImageNet-21K. Most of them are exact duplicates which add no information to the data but waste on storage and compute. In addition 104,000 train\/val leaks were identified by comparing similar images across the train and validation subsets.&#8221;<br \/>\n\u2014Danny Bickson, CEO Visual Layer<\/p><\/blockquote>\n\n<p>&nbsp;<\/p>\n<div class=\"bottom-note\">\n  <span class=\"mod1\">CC-BY-NC Science Surf , accessed 04.08.2026<\/span>\n <\/div>","protected":false},"excerpt":{"rendered":"<p>I think the AI Community is slowly getting at the point where epidemiologists had already been two decades ago, so Ng: \u201cIn many industries where giant data sets simply don\u2019t exist, I think the focus has to shift from big data to good data. Having 50 thoughtfully engineered examples can be sufficient to explain to &hellip; <a href=\"https:\/\/www.wjst.de\/blog\/sciencesurf\/2022\/12\/from-big-data-to-good-data\/\" class=\"more-link\">Continue reading <span class=\"screen-reader-text\">From big data to good data<\/span> <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[3358],"class_list":["post-21381","post","type-post","status-publish","format-standard","hentry","category-computer-software","tag-ai"],"_links":{"self":[{"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/posts\/21381","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/comments?post=21381"}],"version-history":[{"count":4,"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/posts\/21381\/revisions"}],"predecessor-version":[{"id":22971,"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/posts\/21381\/revisions\/22971"}],"wp:attachment":[{"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/media?parent=21381"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/categories?post=21381"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.wjst.de\/blog\/wp-json\/wp\/v2\/tags?post=21381"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}