Format of the input dataset for Google AutoML Natural Language multi-label text classification

1.3k Views Asked by Behzad At 27 July 2025 at 10:14

What should the format of the input dataset be for Google AutoML Natural Language multi-label text classification? I know that for multi-class classification I need a column of text and another column for labels. The labels column include one label per row.

I have multiple labels for each text and I want to do multi-label classification. I tried having one column per label and one-hot encoding but I got this error message: Max 1000 labels supported. Found 9823 labels.

Original Q&A

There are 3 best solutions below

Tim Hong On 28 January 2019 at 22:25

Google AutoML has updated their parser. The following format is fine:

text1, label1, label2, label3,
text1, label1, label2, ,
text1, label1, label2, , ,

At least that worked for me on 27th Jan 2019

Mona Attariyan On 24 October 2018 at 22:59

One column per label is the way to go. If you have less than 1000 labels, you probably have a mistake in your CSV file, where the parser is getting confused and thinks some of the tokens in the text of the example are labels. Please make sure that your text is correctly escaped with quotes around.

Behzad On 25 October 2018 at 21:12

It was very confusing at first but later I managed to find the format in the documentation, which is a CSV file like:

text1, label1, label2 text2, label2 text3, label3, label2, label1

The parser doesn't understand a table with NULL cells saved as a standard CSV file, which is like:

text1, label1, label2, text2, label2,, text3, label3, label2, label1

I had to manually remove extra commas from the CSV file generated by Pandas.

Format of the input dataset for Google AutoML Natural Language multi-label text classification

There are 3 best solutions below

Related Questions in GOOGLE-CLOUD-NL

Related Questions in GOOGLE-NATURAL-LANGUAGE

Related Questions in GOOGLE-CLOUD-AUTOML-NL

Trending Questions

Popular # Hahtags

Popular Questions