ホームページ  >  記事  >  バックエンド開発  >  ETL: テキストから人名を抽出する

ETL: テキストから人名を抽出する

Linda Hamilton
Linda Hamiltonオリジナル
2024-10-08 06:20:30751ブラウズ

Let's say we want to scrape chicagomusiccompass.com.

As you can see, it has several cards, each representing an event. Now, let's check out the next one:

ETL: Extracting a Person

Notice that the name of the event is:


jazmin bean: the traumatic livelihood tour


So now the question is: How do we extract the artist's name from the text?

As a human, I can "easily" tell that jazmin bean is the artist—just check out their wiki page. But writing code to extract that name can get tricky.

We could think, "Hey, anything before the : should be the artist's name," which seems clever, right? It works for this case, but what about this one:


happy hour on the patio: kathryn & chris


Here, the order is flipped. We could keep adding logic to handle different cases, but soon we'll end up with a ton of rules that are fragile and probably won't cover everything.

That’s where Named Entity Recognition (NER) models come in handy. They’re open source and can help us extract names from text. It won’t catch every case, but most of the time, they’ll get us the info we need.

With this approach, the extraction becomes way easier. I'm going with Python because the community around Machine Learning in Python is just unbeatable.


from gliner import GLiNER

model = GLiNER.from_pretrained("urchade/gliner_base")

text = "jazmin bean: the traumatic livelihood tour"
labels = ["person", "bands", "projects"]
entities = model.predict_entities(text, labels)

for entity in entities:
    print(entity["text"], "=>", entity["label"])


Which generates the output:


jazmin bean => person


Now, let’s take a look at that other case:


happy hour on the patio: kathryn & chris


Output:


kathryn => person
chris => person


source-GLiNER

Awesome, right? No more tedious logic to extract names, just use a model. Sure, it won’t cover every possible case, but for my project, this level of flexibility works just fine. If you need more accuracy, you can always:

  • Try a different model
  • Contribute to the existing model
  • Fork the project and tweak it to fit your needs

Conclusion

As a Software Developer, it's highly recommended to stay updated with the tools in the Machine Learning space. Not everything can be solved with just plain programming and logic—some challenges are better tackled using models and statistics.

以上がETL: テキストから人名を抽出するの詳細内容です。詳細については、PHP 中国語 Web サイトの他の関連記事を参照してください。

声明:
この記事の内容はネチズンが自主的に寄稿したものであり、著作権は原著者に帰属します。このサイトは、それに相当する法的責任を負いません。盗作または侵害の疑いのあるコンテンツを見つけた場合は、admin@php.cn までご連絡ください。