Skip to content

Adding new search feature - #69

Open
amckenna41 wants to merge 1 commit into
dahlia:mainfrom
amckenna41:main
Open

amckenna41 wants to merge 1 commit into
dahlia:mainfrom
amckenna41:main

Conversation

@amckenna41

@amckenna41 amckenna41 commented Dec 28, 2025 •

Copy link
Copy Markdown

Hi there 👋

This pull request adds a new and highly useful search functionality to the wikidata package 📦.

Currently in wikidata you can get a Wikidata's entity info via its Entity ID but I was running into issues with actually finding the correct entity ID, having to manually search on the https://www.wikidata.org/w/index.php page and copying and pasting the Entity ID from there. But this new functionality alleviates that, allowing you to search directly within the package for the desired ID, all in one line of code.

For example, I want to find the entity ID for Ireland:
from wikidata.client import Client
client = Client()
client.search("Ireland")

This returns one or more objects of the SearchResult class, which itself contains the attributes entity_id, name, similarity_score and description. The above example will return:

[SearchResult(entity_id='Q27', name='Ireland', similarity_score=100.0%, description='sovereign state in Northwestern Europe'), SearchResult(entity_id='Q22890', name='Ireland', similarity_score=100.0%, description='island in the North Atlantic Ocean'), SearchResult(entity_id='Q28199768', name='Ireland', similarity_score=100.0%, description='family name')]

From this informative output you can find the correct Entity ID that you are looking for and proceed to get that entities data via the client.get function.
entity = client.get('Q27', load=True)

The client.search() function takes the additional optional parameters:

  • limit: the max number of search results to return, default is 10
  • language: language code to search using labels/descriptions in that language
  • entity_type: filter for entity type e.g 'person', 'place', 'organization', filters based on description content
  • likeness_threshold: minimum similarity score (0-1) that results have to match to search input. For example, 0.6 means only results with 60% or higher similarity. By default exact matches (1.0) are returned.
  • verbose: if True, the output results will be pretty printed, default is False.

One of the key additional parameters is the likeness_threshold which allows you to increase or decrease the search space for your search term. Decreasing it will increase the number of search results, vice versa. A default of 1.0 is used meaning only exact search results are returned.

search_results = client.search("Egypt", likeness_threshold=0.6, limit=5)

[SearchResult(entity_id='Q79', name='Egypt', similarity_score=100.0%, description='country in Northeast Africa and Southwest Asia'), SearchResult(entity_id='Q2083973', name='Egypt', similarity_score=100.0%, description='town in Arkansas'), SearchResult(entity_id='Q50868', name='Egyptian', similarity_score=76.9%, description='extinct language spoken in ancient Egypt'), SearchResult(entity_id='Q202311', name='Roman Egypt', similarity_score=62.5%, description='Roman province that encompassed most of modern-day Egypt')]

A search.rst file has been added to the /docs folder as well as a search_test.py test function with a verbose collection of tests. The main README.rst has been updated with some examples of the search functionality.

All tests are passing successfully.

I hope this new feature helps as I've found it fairly useful in one of my current projects that uses the wikidata API. Please provide any comments/feedback if any and happy to discuss further. It would be great if this could be implemented into the package soon for my use in my current project.

Many Thanks!
AJ.

@dahlia dahlia left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hanks for adding this! Search functionality would be a great addition.

I've left some comments—the main suggestion is to use the official wbsearchentities API instead of HTML parsing. Happy to discuss.

Comment thread wikidata/search.py
http_opener = opener

# Build search URL
search_url = f'{search_base_url}w/index.php?search={urllib.parse.quote(query)}'

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This implementation scrapes the HTML search page, which is fragile—any change to Wikidata's page structure will break this.

Wikidata provides an official API for this purpose:

/w/api.php?action=wbsearchentities&search=Ireland&language=en&format=json

The API returns structured JSON with entity IDs, labels, descriptions, and relevance ranking. Would you consider using the official API instead?

Comment thread wikidata/search.py
Comment on lines +10 to +14
if TYPE_CHECKING:
from .client import Client
from urllib.request import OpenerDirector
else:
WIKIDATA_BASE_URL = 'https://www.wikidata.org/'

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

WIKIDATA_BASE_URL is defined only when TYPE_CHECKING is False. This works at runtime, but the intent is unclear. Consider importing it from client.py instead, or moving it outside the conditional block.

Comment thread wikidata/search.py
Comment on lines +154 to +158
if entity_type:
sorted_results = [
result for result in sorted_results
if result.description and entity_type.lower() in result.description.lower()
]

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filtering by substring match on description is unreliable. For example, searching for entity_type="person" would match “personal computer” or “in person.”

The official API has a type parameter (item, property) for this purpose.

Comment thread wikidata/search.py
Comment on lines +414 to +415
except Exception as e:
return [] No newline at end of file

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Catching all exceptions and returning an empty list silently hides errors. Consider either logging the exception or letting it propagate.

Comment thread wikidata/search.py
:param results: List of SearchResult objects to filter
:type results: :class:`List`\\ [:class:`SearchResult`]
:param min_similarity: Minimum similarity threshold (0-100)
:type min_similarity: :class:`float`W

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A small typo:

Suggested change
:type min_similarity: :class:`float`W
:type min_similarity: :class:`float`

Comment thread wikidata/search.py
Comment on lines +50 to +51
opener: Optional['OpenerDirector'] = None,
client: Optional['Client'] = None,

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It accepts both opener and client parameters. Since client already contains opener, consider simplifying to accept only client (or make it a method of Client).

Comment thread wikidata/search.py
Comment on lines +142 to +151
# Calculate similarity scores
for result in parser.results:
result.similarity_score = _calculate_similarity(query, result.name)

# Sort by similarity score (highest first)
sorted_results = sorted(
parser.results,
key=lambda r: r.similarity_score,
reverse=True
)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The likeness_threshold parameter uses difflib.SequenceMatcher to compute string similarity client-side. However, Wikidata's search already returns results ranked by relevance.

Is this additional filtering necessary? It might filter out legitimate results (e.g., “Republic of Ireland” when searching “Ireland”).

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @dahlia for the feedback. I'll look into your comments and get back to you : )

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants