{"id":1380,"date":"2026-06-23T16:30:00","date_gmt":"2026-06-23T16:30:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/"},"modified":"2026-06-23T16:59:27","modified_gmt":"2026-06-23T16:59:27","slug":"i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/","title":{"rendered":"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<p class=\"wp-block-paragraph\"> a big period of time on knowledge preparation for downstream duties. Whether or not it entails knowledge cleansing, dealing with lacking values, function engineering, knowledge preprocessing, or publish processing, this section requires a whole lot of time.<\/p>\n<p class=\"wp-block-paragraph\">So, I used to be engaged on this post-processing activity the place I wanted to create a brand new column in a Pandas DataFrame by extracting values from an current column, based mostly on the information from two different columns. <\/p>\n<p class=\"wp-block-paragraph\">I might have immediately requested an LLM to write down the code (which I normally do) however this time I wished to do it myself. It was early within the morning and I had a contemporary thoughts so I used to be within the temper to deal with some complicated knowledge operations.<\/p>\n<p class=\"wp-block-paragraph\">Here&#8217;s what I needed to do. I had a DataFrame with predicted_categories, pred_category_id, and text_predicted_probs columns.<\/p>\n<p class=\"wp-block-paragraph\">The values within the predicted_categories column are lists of 5 classes in \u201ccategory_id\u201d \u2013 \u201ccategory_description\u201d format. <\/p>\n<p>[&#8216;80814001 &#8211; Freze U\u00e7lar\u0131&#8217;,<br \/>\n &#8216;13003106 &#8211; Freze&#8217;,<br \/>\n &#8216;80805004 &#8211; Sanayi Makineleri&#8217;,<br \/>\n &#8216;13003144 &#8211; Torna Makinesi&#8217;,<br \/>\n &#8216;13003195 &#8211; Kumpas&#8217;]<\/p>\n<p class=\"wp-block-paragraph\">The text_predicted_probs column has the anticipated possibilities of those 5 classes so as. <\/p>\n<p>[0.943, 0.018, 0.008, 0.006, 0.004]<\/p>\n<p class=\"wp-block-paragraph\">Therefore, the primary worth within the text_predicted_probs is the chance of the primary class within the predicted_categories, and so forth.<\/p>\n<p class=\"wp-block-paragraph\">The pred_category_id column reveals the anticipated class id from one other mannequin . What I want is the anticipated chance of the class within the pred_category_id column.<\/p>\n<p class=\"wp-block-paragraph\">I must get the order of the pred_category_id within the predicted_categories column after which take its worth from the test_predicted_probs column.<\/p>\n<p class=\"wp-block-paragraph\">The drawing beneath demonstrates what I wish to obtain:<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/06\/image-274-1024x546.png\" alt=\"\" class=\"wp-image-667992\"\/><\/figure>\n<p class=\"wp-block-paragraph\">On this drawing, I wish to get the chance of class 13003106, which is the second merchandise within the checklist and its corresponding chance worth is 0.018.<\/p>\n<p class=\"wp-block-paragraph\">If we requested Gemini, or one other superior mannequin, we\u2019ll in all probability get the reply in seconds. However, I wished to do it alone first after which ask Gemini.<\/p>\n<p class=\"wp-block-paragraph\">Let\u2019s begin with studying the dataset right into a Pandas DataFrame.<\/p>\n<p>import pandas as pd<\/p>\n<p>outcomes = pd.read_csv(&#8220;prediction_results.csv&#8221;)<\/p>\n<p class=\"wp-block-paragraph\">The values within the predicted_categories column are lists of strings with class ids and class names:<\/p>\n<p>outcomes.loc[0, &#8220;predicted_categories&#8221;]<br \/>\n# output: &#8220;[&#8216;80814001 &#8211; Freze U\u00e7lar\u0131&#8217;, &#8216;13003106 &#8211; Freze&#8217;, &#8216;80805004 &#8211; Sanayi Makineleri&#8217;, &#8216;13003144 &#8211; Torna Makinesi&#8217;, &#8216;13003195 &#8211; Kumpas&#8217;]&#8221;<\/p>\n<p class=\"wp-block-paragraph\">It\u2019s a listing however saved as a string so we first convert it to a listing object utilizing the literal_eval operate within the built-in ast module of Python:<\/p>\n<p>ast.literal_eval(outcomes.loc[0, &#8220;predicted_categories&#8221;])<br \/>\n# output:<br \/>\n[&#8216;80814001 &#8211; Freze U\u00e7lar\u0131&#8217;,<br \/>\n &#8216;13003106 &#8211; Freze&#8217;,<br \/>\n &#8216;80805004 &#8211; Sanayi Makineleri&#8217;,<br \/>\n &#8216;13003144 &#8211; Torna Makinesi&#8217;,<br \/>\n &#8216;13003195 &#8211; Kumpas&#8217;]<\/p>\n<p class=\"wp-block-paragraph\">To extract the class ids, we are able to break up every string on this checklist on the \u201c-\u201d character after which choose the primary half after splitting. Since we&#8217;ve got a listing with 5 classes, we should always do that operation in a listing comprehension as follows:<\/p>\n<p>[category.split(&#8220;-&#8220;)[0].strip()<br \/>\nfor class in ast.literal_eval(outcomes.loc[0, &#8220;predicted_categories&#8221;])]<br \/>\n# output:<br \/>\n[&#8216;80814001&#8217;, &#8216;13003106&#8217;, &#8216;80805004&#8217;, &#8216;13003144&#8217;, &#8216;13003195&#8217;]<\/p>\n<p class=\"wp-block-paragraph\">We\u2019ve completed it for a single worth (i.e. one row). With a purpose to do the identical operation to all the predicted_categories column, we are able to use a listing comprehension. Will probably be a listing comprehension inside one other checklist comprehension (i.e. nested checklist comprehension):<\/p>\n<p>outcomes.loc[:, &#8220;predicted_category_ids&#8221;] = [<br \/>\n    [category.split(&#8220;-&#8220;)[0].strip() for class in ast.literal_eval(predicted_categories)]<br \/>\n    for predicted_categories in outcomes[&#8220;predicted_categories&#8221;]<br \/>\n]<\/p>\n<p class=\"wp-block-paragraph\">We now have class ids extracted from the predicted_categories column:<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/06\/image-299-1024x183.png\" alt=\"\" class=\"wp-image-668608\"\/><\/figure>\n<p class=\"wp-block-paragraph\">The subsequent step is to examine the order of the classes within the predicted class id lists. We&#8217;ll then use this order to extract the anticipated chance of the class.<\/p>\n<p class=\"wp-block-paragraph\">Python checklist object has an index methodology, which returns the index (i.e. order) of the merchandise within the checklist.<\/p>\n<p>outcomes.loc[0, &#8220;predicted_category_ids&#8221;]<br \/>\n# output:<br \/>\n[&#8216;80814001&#8217;, &#8216;13003106&#8217;, &#8216;80805004&#8217;, &#8216;13003144&#8217;, &#8216;13003195&#8217;]<\/p>\n<p>outcomes.loc[0, &#8220;predicted_category_ids&#8221;].index(&#8220;13003106&#8221;)<br \/>\n# output:<br \/>\n2<\/p>\n<p class=\"wp-block-paragraph\">As soon as I discover the index of a predicted class id, I can use it to get the chance of this class id from the text_predicted_probs column:<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/06\/image-298-1024x251.png\" alt=\"\" class=\"wp-image-668607\"\/><\/figure>\n<p class=\"wp-block-paragraph\">What we have to do:<\/p>\n<p>Get the index of pred_category_id within the predicted_category_ids<\/p>\n<p>Use this index to extract the related worth from text_predicted_probs<\/p>\n<p class=\"wp-block-paragraph\">These steps could be completed in a single operation by zipping these three columns. Let\u2019s take a look at it on the primary row:<\/p>\n<p>for i, j, ok in zip(outcomes[&#8220;pred_category_id&#8221;][:1], outcomes[&#8220;predicted_category_ids&#8221;][:1], outcomes[&#8220;text_predicted_probs&#8221;][:1]):<br \/>\n    print(j.index(str(i))) # get the index of pred_category_id in predicted_category_ids<br \/>\n    print(ast.literal_eval(ok)[j.index(str(i))]) # get the worth at this index in text_predicted_probs<\/p>\n<p># output:<br \/>\n0<br \/>\n0.943<\/p>\n<p class=\"wp-block-paragraph\">We are able to affirm the output within the screenshot above. The pred_category_id is 80814001, which is the primary merchandise (i.e. index = 0) within the predicted_category_ids and the primary chance worth is 0.943.<\/p>\n<p class=\"wp-block-paragraph\">The lists within the text_predicted_probs column are saved as string as effectively so we apply the literal_eval operate to transform them to a listing object.<\/p>\n<p class=\"wp-block-paragraph\">We are able to do that operation on all the DataFrame utilizing a listing comprehension as observe:<\/p>\n<p>outcomes.loc[:, &#8220;pred_category_prob&#8221;] = [<br \/>\n    float(ast.literal_eval(k)[j.index(str(i))]) if str(i) in j else 0<br \/>\n    for i, j, ok in zip(outcomes[&#8220;pred_category_id&#8221;], outcomes[&#8220;predicted_category_ids&#8221;], outcomes[&#8220;text_predicted_probs&#8221;])<br \/>\n]<\/p>\n<p class=\"wp-block-paragraph\">We mainly transformed the for loop within the earlier code block to a listing comprehension. I\u2019ve solely added a examine \u201cif str(i) in j else 0\u201d as a result of there are some rows with a pred_category_id that doesn&#8217;t exist in predicted_category_ids. In such circumstances, we set the chance worth to 0.<\/p>\n<p class=\"wp-block-paragraph\">The duty is now accomplished as we are able to see beneath:<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/contributor.insightmediagroup.io\/wp-content\/uploads\/2026\/06\/image-300-1024x214.png\" alt=\"\" class=\"wp-image-668609\"\/><\/figure>\n<h2 class=\"wp-block-heading\">How did Gemini do that?<\/h2>\n<p class=\"wp-block-paragraph\">After working about an hour (together with breaks) I accomplished the duty. Then, I wished to ask Gemini to see how it will deal with the identical activity.<\/p>\n<p class=\"wp-block-paragraph\">I wrote a immediate explaining my aim and included a small pattern of the information.<\/p>\n<p class=\"wp-block-paragraph\">Here&#8217;s what Gemini produced on its first try:<\/p>\n<p>import pandas as pd<\/p>\n<p># 1. Create the pattern DataFrame<br \/>\nknowledge = {<br \/>\n    &#8216;predicted_categories&#8217;: [<br \/>\n        [&#8216;80814001 &#8211; Freze U\u00e7lar\u0131&#8217;, &#8216;13003106 &#8211; Freze&#8217;, &#8216;80805004 &#8211; Sanayi Makineleri&#8217;, &#8216;13003144 &#8211; Torna Makinesi&#8217;, &#8216;13003195 &#8211; Kumpas&#8217;]<br \/>\n    ],<br \/>\n    &#8216;pred_category_id&#8217;: [80814001],<br \/>\n    &#8216;text_predicted_probs&#8217;: [<br \/>\n        [0.943, 0.018, 0.008, 0.006, 0.004]<br \/>\n    ]<br \/>\n}<\/p>\n<p>df = pd.DataFrame(knowledge)<\/p>\n<p># 2. Outline a operate to extract the chance matching the class ID<br \/>\ndef get_matching_prob(row):<br \/>\n    # Convert ID to string for matching<br \/>\n    target_id_str = str(row[&#8216;pred_category_id&#8217;])<\/p>\n<p>    # Iterate by the classes to seek out the matching index<br \/>\n    for index, class in enumerate(row[&#8216;predicted_categories&#8217;]):<br \/>\n        if class.startswith(target_id_str):<br \/>\n            # Return the corresponding chance from the identical index<br \/>\n            return row[&#8216;text_predicted_probs&#8217;][index]<\/p>\n<p>    # Return None (or 0) if no match is discovered to stop the code from crashing<br \/>\n    return None<\/p>\n<p># 3. Create the brand new column<br \/>\ndf[&#8216;pred_category_prob&#8217;] = df.apply(get_matching_prob, axis=1)<\/p>\n<p># Show the end result<br \/>\nprint(df[[&#8216;pred_category_id&#8217;, &#8216;pred_category_prob&#8217;]])<\/p>\n<p class=\"wp-block-paragraph\">It first created a pattern DataFrame utilizing the information I supplied within the immediate. Then, it wrote a operate that does precisely what I defined.<\/p>\n<p class=\"wp-block-paragraph\">It took me about an hour to write down the code myself. Gemini accomplished the very same activity in seconds based mostly on a immediate that took me just a few minutes to write down. <\/p>\n<p class=\"wp-block-paragraph\">Nonetheless, there&#8217;s a minor subject in Gemini\u2019s code. It makes use of the apply operate, which isn&#8217;t vectorized. In case of working with massive datasets, utilizing the apply operate (mainly a for loop) can improve the run time. <\/p>\n<p class=\"wp-block-paragraph\">However the resolution is easy. I requested Gemini to write down a vectorized model of it and here&#8217;s what I acquired:<\/p>\n<p>import pandas as pd<\/p>\n<p># Pattern DataFrame<br \/>\nknowledge = {<br \/>\n    &#8216;predicted_categories&#8217;: [<br \/>\n        [&#8216;80814001 &#8211; Freze U\u00e7lar\u0131&#8217;, &#8216;13003106 &#8211; Freze&#8217;, &#8216;80805004 &#8211; Sanayi Makineleri&#8217;],<br \/>\n        [&#8216;13003144 &#8211; Torna Makinesi&#8217;, &#8216;13003195 &#8211; Kumpas&#8217;]<br \/>\n    ],<br \/>\n    &#8216;pred_category_id&#8217;: [80814001, 13003195],<br \/>\n    &#8216;text_predicted_probs&#8217;: [<br \/>\n        [0.943, 0.018, 0.008],<br \/>\n        [0.6, 0.4]<br \/>\n    ]<br \/>\n}<br \/>\ndf = pd.DataFrame(knowledge)<\/p>\n<p># 1. Explode each checklist columns concurrently to keep up alignment between class and chance<br \/>\ndf_exploded = df.explode([&#8216;predicted_categories&#8217;, &#8216;text_predicted_probs&#8217;])<\/p>\n<p># 2. Extract the numeric ID from the class string utilizing vectorized regex<br \/>\ndf_exploded[&#8216;extracted_id&#8217;] = df_exploded[&#8216;predicted_categories&#8217;].str.extract(r&#8217;^(d+)&#8217;).astype(float)<\/p>\n<p># 3. Filter for rows the place the extracted ID matches the goal &#8216;pred_category_id&#8217;<br \/>\nmatched = df_exploded[df_exploded[&#8216;extracted_id&#8217;] == df_exploded[&#8216;pred_category_id&#8217;]]<\/p>\n<p># 4. Dedup the index (security internet in case an ID seems twice inside the identical checklist)<br \/>\nmatched = matched[~matched.index.duplicated(keep=&#8217;first&#8217;)]<\/p>\n<p># 5. Map the extracted chance column again to the unique DataFrame utilizing the index<br \/>\ndf[&#8216;pred_category_prob&#8217;] = matched[&#8216;text_predicted_probs&#8217;]<\/p>\n<p>df<\/p>\n<p class=\"wp-block-paragraph\">The second resolution was completely advantageous and seemed less complicated than the code I wrote.<\/p>\n<p class=\"wp-block-paragraph\">So, I spent about an hour on a activity that an LLM might have accomplished in lower than 5 minutes. Nonetheless, if I didn\u2019t know the way Pandas labored, I&#8217;d have accepted the primary resolution, which was not the optimum one. It&#8217;s a good instance of how LLMs can improve productiveness, however provided that you really know what you\u2019re doing.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/towardsdatascience.com\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>a big period of time on knowledge preparation for downstream duties. Whether or not it entails knowledge cleansing, dealing with lacking values, function engineering, knowledge preprocessing, or publish processing, this section requires a whole lot of time. So, I used to be engaged on this post-processing activity the place I wanted to create a brand [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1382,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[7],"tags":[160,287,1125,1829,1828,1830],"class_list":["post-1380","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-data-science-mlops","tag-data","tag-gemini","tag-hour","tag-preprocessing","tag-spent","tag-task"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini - Future News 24<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini - Future News 24\" \/>\n<meta property=\"og:description\" content=\"a big period of time on knowledge preparation for downstream duties. Whether or not it entails knowledge cleansing, dealing with lacking values, function engineering, knowledge preprocessing, or publish processing, this section requires a whole lot of time. So, I used to be engaged on this post-processing activity the place I wanted to create a brand [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-23T16:30:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-23T16:59:27+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini\",\"datePublished\":\"2026-06-23T16:30:00+00:00\",\"dateModified\":\"2026-06-23T16:59:27+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/\"},\"wordCount\":1559,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg\",\"keywords\":[\"data\",\"Gemini\",\"hour\",\"Preprocessing\",\"Spent\",\"Task\"],\"articleSection\":[\"Data Science &amp; MLOps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/\",\"name\":\"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg\",\"datePublished\":\"2026-06-23T16:30:00+00:00\",\"dateModified\":\"2026-06-23T16:59:27+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#primaryimage\",\"url\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg\",\"contentUrl\":\"https:\\\/\\\/towardsdatascience.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/06\\\/23\\\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini - Future News 24","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/","og_locale":"en_US","og_type":"article","og_title":"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini - Future News 24","og_description":"a big period of time on knowledge preparation for downstream duties. Whether or not it entails knowledge cleansing, dealing with lacking values, function engineering, knowledge preprocessing, or publish processing, this section requires a whole lot of time. So, I used to be engaged on this post-processing activity the place I wanted to create a brand [&hellip;]","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/","og_site_name":"Future News 24","article_published_time":"2026-06-23T16:30:00+00:00","article_modified_time":"2026-06-23T16:59:27+00:00","og_image":[{"url":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg","twitter_misc":{"Written by":"Future News 24","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini","datePublished":"2026-06-23T16:30:00+00:00","dateModified":"2026-06-23T16:59:27+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/"},"wordCount":1559,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#primaryimage"},"thumbnailUrl":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg","keywords":["data","Gemini","hour","Preprocessing","Spent","Task"],"articleSection":["Data Science &amp; MLOps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/","name":"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#primaryimage"},"thumbnailUrl":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg","datePublished":"2026-06-23T16:30:00+00:00","dateModified":"2026-06-23T16:59:27+00:00","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#primaryimage","url":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg","contentUrl":"https:\/\/towardsdatascience.com\/wp-content\/uploads\/2026\/06\/national-institute-of-allergy-and-infectious-diseases-oc12eprOeoI-unsplash-scaled-1.jpg"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/06\/23\/i-spent-an-hour-on-a-data-preprocessing-task-before-asking-gemini\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"I Spent an Hour on a Information Preprocessing Job Earlier than Asking Gemini"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1380","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=1380"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1380\/revisions"}],"predecessor-version":[{"id":1381,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/1380\/revisions\/1381"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/1382"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=1380"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=1380"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=1380"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}