Repository navigation
Adding new search feature - #69
amckenna41 wants to merge 1 commit into
Conversation
dahlia
left a comment
There was a problem hiding this comment.
hanks for adding this! Search functionality would be a great addition.
I've left some comments—the main suggestion is to use the official wbsearchentities API instead of HTML parsing. Happy to discuss.
| http_opener = opener | ||
|
|
||
| # Build search URL | ||
| search_url = f'{search_base_url}w/index.php?search={urllib.parse.quote(query)}' |
There was a problem hiding this comment.
This implementation scrapes the HTML search page, which is fragile—any change to Wikidata's page structure will break this.
Wikidata provides an official API for this purpose:
/w/api.php?action=wbsearchentities&search=Ireland&language=en&format=json
The API returns structured JSON with entity IDs, labels, descriptions, and relevance ranking. Would you consider using the official API instead?
| if TYPE_CHECKING: | ||
| from .client import Client | ||
| from urllib.request import OpenerDirector | ||
| else: | ||
| WIKIDATA_BASE_URL = 'https://www.wikidata.org/' |
There was a problem hiding this comment.
WIKIDATA_BASE_URL is defined only when TYPE_CHECKING is False. This works at runtime, but the intent is unclear. Consider importing it from client.py instead, or moving it outside the conditional block.
| if entity_type: | ||
| sorted_results = [ | ||
| result for result in sorted_results | ||
| if result.description and entity_type.lower() in result.description.lower() | ||
| ] |
There was a problem hiding this comment.
Filtering by substring match on description is unreliable. For example, searching for entity_type="person" would match “personal computer” or “in person.”
The official API has a type parameter (item, property) for this purpose.
| except Exception as e: | ||
| return [] No newline at end of file |
There was a problem hiding this comment.
Catching all exceptions and returning an empty list silently hides errors. Consider either logging the exception or letting it propagate.
| :param results: List of SearchResult objects to filter | ||
| :type results: :class:`List`\\ [:class:`SearchResult`] | ||
| :param min_similarity: Minimum similarity threshold (0-100) | ||
| :type min_similarity: :class:`float`W |
There was a problem hiding this comment.
A small typo:
| :type min_similarity: :class:`float`W | |
| :type min_similarity: :class:`float` |
| opener: Optional['OpenerDirector'] = None, | ||
| client: Optional['Client'] = None, |
There was a problem hiding this comment.
It accepts both opener and client parameters. Since client already contains opener, consider simplifying to accept only client (or make it a method of Client).
| # Calculate similarity scores | ||
| for result in parser.results: | ||
| result.similarity_score = _calculate_similarity(query, result.name) | ||
|
|
||
| # Sort by similarity score (highest first) | ||
| sorted_results = sorted( | ||
| parser.results, | ||
| key=lambda r: r.similarity_score, | ||
| reverse=True | ||
| ) |
There was a problem hiding this comment.
The likeness_threshold parameter uses difflib.SequenceMatcher to compute string similarity client-side. However, Wikidata's search already returns results ranked by relevance.
Is this additional filtering necessary? It might filter out legitimate results (e.g., “Republic of Ireland” when searching “Ireland”).
There was a problem hiding this comment.
Thanks @dahlia for the feedback. I'll look into your comments and get back to you : )
Hi there 👋
This pull request adds a new and highly useful search functionality to the wikidata package 📦.
Currently in wikidata you can get a Wikidata's entity info via its Entity ID but I was running into issues with actually finding the correct entity ID, having to manually search on the https://www.wikidata.org/w/index.php page and copying and pasting the Entity ID from there. But this new functionality alleviates that, allowing you to search directly within the package for the desired ID, all in one line of code.
For example, I want to find the entity ID for Ireland:
from wikidata.client import Clientclient = Client()client.search("Ireland")This returns one or more objects of the SearchResult class, which itself contains the attributes entity_id, name, similarity_score and description. The above example will return:
[SearchResult(entity_id='Q27', name='Ireland', similarity_score=100.0%, description='sovereign state in Northwestern Europe'), SearchResult(entity_id='Q22890', name='Ireland', similarity_score=100.0%, description='island in the North Atlantic Ocean'), SearchResult(entity_id='Q28199768', name='Ireland', similarity_score=100.0%, description='family name')]From this informative output you can find the correct Entity ID that you are looking for and proceed to get that entities data via the
client.getfunction.entity = client.get('Q27', load=True)The
client.search()function takes the additional optional parameters:One of the key additional parameters is the likeness_threshold which allows you to increase or decrease the search space for your search term. Decreasing it will increase the number of search results, vice versa. A default of 1.0 is used meaning only exact search results are returned.
search_results = client.search("Egypt", likeness_threshold=0.6, limit=5)[SearchResult(entity_id='Q79', name='Egypt', similarity_score=100.0%, description='country in Northeast Africa and Southwest Asia'), SearchResult(entity_id='Q2083973', name='Egypt', similarity_score=100.0%, description='town in Arkansas'), SearchResult(entity_id='Q50868', name='Egyptian', similarity_score=76.9%, description='extinct language spoken in ancient Egypt'), SearchResult(entity_id='Q202311', name='Roman Egypt', similarity_score=62.5%, description='Roman province that encompassed most of modern-day Egypt')]A
search.rstfile has been added to the /docs folder as well as asearch_test.pytest function with a verbose collection of tests. The mainREADME.rsthas been updated with some examples of the search functionality.All tests are passing successfully.
I hope this new feature helps as I've found it fairly useful in one of my current projects that uses the wikidata API. Please provide any comments/feedback if any and happy to discuss further. It would be great if this could be implemented into the package soon for my use in my current project.
Many Thanks!
AJ.