The Wayback Machine - https://web.archive.org/web/20240926114441/https://www.geeksforgeeks.org/python-check-url-string/
Open In App

Python | Check for URL in a String

Last Updated : 25 Apr, 2023
Summarize
Comments
Improve
Suggest changes
Like Article
Like
Save
Share
Report
News Follow

Prerequisite: Pattern matching with Regular Expression In this article, we will need to accept a string and we need to check if the string contains any URL in it. If the URL is present in the string, we will say URL’s been found or not and print the respective URL present in the string. We will use the concept of Regular Expression of Python to solve the problem.

Examples:

Input : string = 'My Profile: 
https://auth.geeksforgeeks.org/user/Chinmoy%20Lenka/articles 
in the portal of https://www.geeksforgeeks.org/'

Output : URLs :  ['https://auth.geeksforgeeks.org/user/Chinmoy%20Lenka/articles',
'https://www.geeksforgeeks.org/']

Input : string = 'I am a blogger at https://geeksforgeeks.org'
Output : URL :  ['https://geeksforgeeks.org']

To find the URLs in a given string we have used the findall() function from the regular expression module of Python. This return all non-overlapping matches of pattern in string, as a list of strings. The string is scanned left to right, and matches are returned in the order found. 

Python3




# Python code to find the URL from an input string
# Using the regular expression
import re
 
 
def Find(string):
 
    # findall() has been used
    # with valid conditions for urls in string
    regex = r"(?i)\b((?:https?://|www\d{0,3}[.]|[a-z0-9.\-]+[.][a-z]{2,4}/)(?:[^\s()<>]+|\(([^\s()<>]+|(\([^\s()<>]+\)))*\))+(?:\(([^\s()<>]+|(\([^\s()<>]+\)))*\)|[^\s`!()\[\]{};:'\".,<>?«»“”‘’]))"
    url = re.findall(regex, string)
    return [x[0] for x in url]
 
 
# Driver Code
print("Urls: ", Find(string))


Output

Urls:  ['https://auth.geeksforgeeks.org/user/Chinmoy%20Lenka/articles', 'https://www.geeksforgeeks.org/']

Time complexity: O(n) where n is the length of the input string, as the findall function of the re library iterates through the entire string once to find the URLs.
Auxiliary space: O(n), where n is the length of the input string, as the function Find stores all the URLs found in the input string into a list which is returned at the end.

Method #2: Using startswith() method

Python3




# Python code to find the URL from an input string
 
def Find(string):
    x=string.split()
    res=[]
    for i in x:
        if i.startswith("https:") or i.startswith("http:"):
            res.append(i)
    return res
             
# Driver Code
print("Urls: ", Find(string))


Output

Urls:  ['https://auth.geeksforgeeks.org/user/Chinmoy%20Lenka/articles', 'https://www.geeksforgeeks.org/']

Time Complexity: O(n), where n is the number of words in the input string.
Auxiliary Space: O(n), where n is the number of URLs found in the input string. The auxiliary space is used to store the found URLs in the res list.

Method #3 : Using find() method

Python3




# Python code to find the URL from an input string
 
def Find(string):
    x=string.split()
    res=[]
    for i in x:
        if i.find("https:")==0 or i.find("http:")==0:
            res.append(i)
    return res
             
# Driver Code
print("Urls: ", Find(string))


Output

Urls:  ['https://auth.geeksforgeeks.org/user/Chinmoy%20Lenka/articles', 'https://www.geeksforgeeks.org/']

Time Complexity : O(N)
Auxiliary Space : O(N)

METHOD 4:Using the urlparse() function from urllib.parse

APPROACH:

The urlparse() function from the urllib.parse module in Python can be used to extract various components of a URL, such as the scheme, netloc, path, query, and fragment

ALGORITHM:

1.Import the required modules – urllib.parse for urlparse() and split() methods from Python’s inbuilt module string
2.Initialize the input string.
3.Split the input string into individual words using the split() method.
4.Initialize an empty list urls to store the extracted URLs.
5.Iterate through each word in the words list.
6.Use the urlparse() function to extract the scheme and netloc components of the URL.
7.Check if both scheme and netloc are present in the URL, indicating that it is a valid URL.
8.If the URL is valid, add it to the urls list.
9.Print the final list of extracted URLs.

Python3




from urllib.parse import urlparse
 
 
# Split the string into words
words = string.split()
 
# Extract URLs from the words using urlparse()
urls = []
for word in words:
    parsed = urlparse(word)
    if parsed.scheme and parsed.netloc:
        urls.append(word)
 
# Print the extracted URLs
print("URLs:", urls)


Output

URLs: ['https://auth.geeksforgeeks.org/user/Chinmoy%20Lenka/articles', 'https://www.geeksforgeeks.org/']

The time complexity of this approach is O(n), where n is the length of the input string

The space complexity of this approach is O(n + k), where n is the length of the input string and k is the number of extracted URLs..

METHOD 5: Using reduce():

Algorithm :

  1. Define a function merge_url_lists(url_list1, url_list2) that takes two lists of URLs and returns their concatenation.
  2. Define a function find_urls_in_string(string) that takes a string as input and returns a list of URLs found in the string.
  3. Define two input strings, string1 and string2, and put them into a list called string_list.
  4. Call map(find_urls_in_string, string_list) to generate a list of lists of URLs found in each string.
  5. Call reduce(merge_url_lists, map(find_urls_in_string, string_list)) to concatenate all the lists of URLs found into a single list.
  6. Print the resulting list of URLs.

Python3




from functools import reduce
 
def merge_url_lists(url_list1, url_list2):
    return url_list1 + url_list2
 
def find_urls_in_string(string):
    x = string.split()
    return [i for i in x if i.find("https:")==0 or i.find("http:")==0]
 
string2 = 'Some more text without URLs'
 
string_list = [string1, string2]
 
url_list = reduce(merge_url_lists, map(find_urls_in_string, string_list))
 
print("Urls:", url_list)
#This code is contributed by Rayudu.


Output

Urls: ['https://auth.geeksforgeeks.org/user/Chinmoy%20Lenka/articles', 'https://www.geeksforgeeks.org/']

The time complexity: O(n*m), where n is the number of strings in the string_list list and m is the maximum number of words in a string. This is because the find_urls_in_string() function needs to split each string into words and check each word for the presence of a URL.

The space complexity : O(n*m), because it stores all the words in all the strings in memory, as well as all the URLs found.



Similar Reads

How to Add URL Parameters to Django Template URL Tag?
When building a Django application, it's common to construct URLs dynamically, especially when dealing with views that require specific parameters. The Django template language provides the {% url %} template tag to create URLs by reversing the URL patterns defined in the urls.py file. This tag becomes especially useful when we need to pass additio
4 min read
Comparing path() and url() (Deprecated) in Django for URL Routing
When building web applications with Django, URL routing is a fundamental concept that allows us to direct incoming HTTP requests to the appropriate view function or class. Django provides two primary functions for defining URL patterns: path() and re_path() (formerly url()). Although both are used for routing, they serve slightly different purposes
8 min read
Building CLI to check status of URL using Python
In this article, we will build a CLI(command-line interface) program to verify the status of a URL using Python. The python CLI takes one or more URLs as arguments and checks whether the URL is accessible (or)not. Stepwise ImplementationStep 1: Setting up files and Installing requirements First, create a directory named "urlcheck" and create a new
4 min read
Check if an URL is valid or not using Regular Expression
Given a URL as a character string str of size N.The task is to check if the given URL is valid or not.Examples : Input : str = "https://www.geeksforgeeks.org/" Output : Yes Explanation : The above URL is a valid URL.Input : str = "https:// www.geeksforgeeks.org/" Output : No Explanation : Note that there is a space after https://, hence the URL is
4 min read
URL Shorteners and its API in Python | Set-2
Prerequisite: URL Shorteners and its API | Set-1 Let's discuss few more URL shorteners using the same pyshorteners module. Bitly URL Shortener: We have seen Bitly API implementation. Now let's see how this module can be used to shorten a URL using Bitly Services. Code to Shorten URL: from pyshorteners import Shortener ACCESS_TOKEN = '82e8156..e4e12
2 min read
Python | Extract URL from HTML using lxml
Link extraction is a very common task when dealing with the HTML parsing. For every general web crawler that's the most important function to perform. Out of all the Python libraries present out there, lxml is one of the best to work with. As explained in this article, lxml provides a number of helper function in order to extract the links. lxml in
4 min read
Django URL patterns | Python
Prerequisites: Views in Django In Django, views are Python functions which take a URL request as parameter and return an HTTP response or throw an exception like 404. Each view needs to be mapped to a corresponding URL pattern. This is done via a Python module called URLConf(URL configuration) Let the project name be myProject. The Python module to
2 min read
Python | Sorting URL on basis of Top Level Domain
Given a list of URL, the task is to sort the URL in the list based on the top-level domain. A top-level domain (TLD) is one of the domains at the highest level in the hierarchical Domain Name System of the Internet. Example - org, com, edu. This is mostly used in a case where we have to scrap the pages and sort URL according to top-level domain. It
3 min read
response.url - Python requests
response.url returns the URL of the response. It will show the main url which has returned the content, after all redirections, if done. Python requests are generally used to fetch the content from a particular resource URI. Whenever we make a request to a specified URI through Python, it returns a response object. Now, this response object would b
2 min read
Python IMDbPY – Getting cover URL of the series
In this article we will see how we can get the cover url of the series, cover url is basically a url which when open show the cover image of the series, cover image is the front or we can say the poster of the series. In order to get this we have to do the following - 1. Get the series details with the help of get_movie method 2. As this object wil
2 min read
Python Tweepy – Getting the URL of a user
In this article we will see how we can get the URL of a user. The url attribute is a URL provided by the user in association with their profile. The URL attribute is optional and is Nullable. Identifying the URL in the GUI : In the above mentioned profile the url is : geeksforgeeks.org In order to get the URL we have to do the following : Identify
2 min read
Requesting a URL from a local File in Python
Making requests over the internet is a common operation performed by most automated web applications. Whether a web scraper or a visitor tracker, such operations are performed by any program that makes requests over the internet. In this article, you will learn how to request a URL from a local File using Python. Requesting a URL from a local File
4 min read
Parsing and Processing URL using Python - Regex
Prerequisite: Regular Expression in Python URL or Uniform Resource Locator consists of many information parts, such as the domain name, path, port number etc. Any URL can be processed and parsed using Regular Expression. So for using Regular Expression we have to use re library in Python. Example: URL: https://www.geeksforgeeks.org/courses When we
3 min read
Build an Application to extract URL and Metadata from a PDF using Python
The PDF (Portable Document Format) is the most common use platform-independent file format developed by Adobe to present documents. There are lots of PDF-related packages for Python, one of them is the pdfx module. The pdfx module is used to extract URL, MetaData, and Plain text from a given PDF or PDF URL. Features: Extract references and metadata
3 min read
Launch Website URL shortcut using Python
In this article, we are going to launch favorite websites using shortcuts, for this, we will use Python's sqlite3 and webbrowser modules to launch your favorite websites using shortcuts. Both sqlite3 and webbrowser are a part of the python standard library, so we don't need to install anything separately. The best part about this is that because we
3 min read
How to Open URL in Firefox Browser from Python Application?
In this article, we'll look at how to use a Python application to access a URL in the Firefox browser. To do so, we'll use the webbrowser Python module. We don't have to install it because it comes pre-installed. There are also a variety of browsers pre-defined in this module, and for this article, we'll be utilizing Firefox. The webbrowser module
2 min read
How to download an image from a URL in Python
Downloading content from its URL is a common task that Web Scrapers or online trackers perform. These URLs or Uniform Resource Locators can contain the web address (or local address) of a webpage, website, image, text document, container files, and many other online resources. It is quite easy to download and store content from files on the interne
3 min read
Python Pyramid - Url Routing
URL routing is a fundamental aspect of web application development, as it defines how different URLs in your application are mapped to specific views or functions. In Python Pyramid, a powerful web framework, URL routing is made simple and flexible. In this article, we will explore the concept of URL routing in Python Pyramid, with practical exampl
5 min read
URL Shorteners and its API in Python | Set-1
URL Shortener, as the name suggests, is a service to help to reduce the length of the URL so that it can be shared easily on platforms like Twitter, where number of characters is an issue. There are so many URL Shorteners available in the market today, that will definitely help you out to solve the purpose. We will be discussing API implementation
5 min read
How To Fix “Connectionerror : Max retries exceeded with URL” In Python?
When the Python script fails to establish a connection with a web resource after a certain number of attempts, the error 'Connection error: Max retries exceeded with URL' occurs. It's often caused by issues like network problems, server unavailability, DNS errors, or request timeouts. In this article, we will see some concepts related to this error
3 min read
Fetch JSON URL Data and Store in Excel using Python
In this article, we will learn how to fetch the JSON data from a URL using Python, parse it, and store it in an Excel file. We will use the Requests library to fetch the JSON data and Pandas to handle the data manipulation and export it to Excel. Fetch JSON data from URL and store it in an Excel file in PythonThe process of fetching JSON data from
3 min read
String slicing in Python to check if a string can become empty by recursive deletion
Given a string “str” and another string “sub_str”. We are allowed to delete “sub_str” from “str” any number of times. It is also given that the “sub_str” appears only once at a time. The task is to find if “str” can become empty by removing “sub_str” again and again. Examples: Input : str = "GEEGEEKSKS", sub_str = "GEEKS" Output : Yes Explanation :
2 min read
url - Django Template Tag
A Django template is a text document or a Python string marked-up using the Django template language. Django being a powerful Batteries included framework provides convenience to rendering data in a template. Django templates not only allow passing data from view to template, but also provides some limited features of programming such as variables,
3 min read
URL fields in serializers - Django REST Framework
In Django REST Framework the very concept of Serializing is to convert DB data to a datatype that can be used by javascript. Every serializer comes with some fields (entries) which are going to be processed. For example if you have a class with name Employee and its fields as Employee_id, Employee_name, is_admin, etc. Then, you would need AutoField
5 min read
Pafy - Getting URL of Stream
In this article we will see how we can get url of the given youtube video stream in pafy. Pafy is a python library to download YouTube content and retrieve metadata. Pafy object is the object which contains all the information about the given video. Stream is basically available resolution of video is available on youtube. URL stands for Uniform Re
2 min read
Pafy - Getting https URL of Stream
In this article we will see how we can get secured url of the given youtube video stream in pafy. Pafy is a python library to download YouTube content and retrieve metadata. Pafy object is the object which contains all the information about the given video. Stream is basically available resolution of video is available on youtube. URL stands for Un
2 min read
Pafy - Getting Watch URL for Each Item of Playlist
In this article we will see how we can get the watch URLfrom playlist items in pafy. Pafy is a python library to download YouTube content and retrieve metadata. Pafy object is the object which contains all the information about the given video. A playlist in YouTube is a list, or group, of videos that plays in order, one video after the other. Watc
2 min read
PYGLET – Creating URL Location
In this article, we will see how we can create a URL location object in PYGLET module in python. Pyglet is easy to use but powerful library for developing visually rich GUI applications like games, multimedia, etc. A window is a "heavyweight" object occupying operating system resources. Windows may appear as floating regions or can be set to fill a
2 min read
How to open an image from the URL in PIL?
In this article, we will learn How to open an image from the URL using the PIL module in python. For the opening of the image from a URL in Python, we need two Packages urllib and Pillow(PIL). Approach:Install the required libraries and then import them. To install use the following commands:pip install pillowCopy the URL of any image.Write URL wit
1 min read
Redirecting to URL in Flask
Flask is a backend server that is built entirely using Python. It is a framework that consists of Python packages and modules. It is lightweight which makes developing backend applications quicker with its features. In this article, we will learn to redirect a URL in the Flask web application. Redirect to a URL in Flask A redirect is used in the Fl
3 min read