One of my post (http://hengrui-li.blogspot.com/2011/08/php-copy-on-write-how-php-manages.html) discusses PHP's copy on write mechanism and explain why passing by value won't cost more memory in most cases.
Actually, pass by reference in PHP is considered as a bad practice, even from the perspective of performance. Today i found these two very valuable posts about PHP's reference mechanism. It definitely worths a read:
1. http://schlueters.de/blog/archives/125-Do-not-use-PHP-references.html
2. http://schlueters.de/blog/archives/141-References-and-foreach.html
Both posts from the same guy and his English is much better than mine. Anyway, here quotes from the summary of his post: "Do not use references for performance". And he explains the reason very well.
We should try to avoid using pass by reference in php. The reason is quite simple and common: when we have a lot of references in our system, changing one will change another and it will be hard for us to track what happened.
So passing by reference for performance is No, what if we want to return multiple value from a function? We can do that in other ways, for example, we can return an array from the function, or we can pass a parameter object into the function.
What if sometimes a function is defined and we don't want to change its return type, and we don't want to build parameter object? This is almost like saying 'I just want to use reference'. Well, honestly, sometimes i use reference as well, for convenience and laziness. When i was tempted to use reference, i always check if this prerequisite holds true:
it is only in a private method of a class. That means, the method using pass by reference should be hidden within the class. It should not be exposed to others. It must be private only(no protected, no public).
Well, even that, avoid using reference is still a generic rule and we should respect it.
Thursday, August 18, 2011
Wednesday, August 17, 2011
UTF-8, multibyte functions in php web application
Output text/string in UTF-8 encoding
There are several ways to tell browser how to encode a page. One way is to specify a meta tag in html:
<head>
<meta http-equiv="Content-Type" content="text/html; charset=utf-8" />
This approach is simple and easy. But it does contain some disadvantages. One issue is, most browsers will have to start re-parsing the document after reaching the meta tag, because they may have already parsed the document with incorrect encoding. This may cause a delay in page rendering. Due to this, it is better that we output the UTF-8 header with PHP.
Output UTF-8 header with PHP:
We use PHP's header function header("Content-Type:text/html;charset=utf-8");
The method is safe enough to ensure the page is encoded with UTF-8. However, it is obviously not as convenient as the way of specifying meta tag of http-equiv, which can be done in a layout template or a header file and then is included in other pages.
We still have the third solution. Assuming we are using Apache web server(i believe most PHP apps run on apache), we can specify the charset in a specific .htaccess file.
Specify chartset in .htaccess file:
AddDefaultCharset utf-8
In this way we send regular http utf-8 head through web server configuration.
PHP multibyte string functions
We know that PHP provides a set of multibyte string functions, which are prefixed with 'mb_'. There are some interesting things about them. Let's take mb_substr for example. I'm using PHP 5.3.3 with the default php.ini configuration. To do the test, nothing could be better than using my native language: Chinese.
<?php
header("Content-Type:text/html;charset=utf-8");
$str = '我爱编程';
echo substr($str, 2, 2), '<br>';
echo mb_substr($str, 2, 2), '<br>';
?>
Ok, let me explain. The first line header("Content-Type:text/html;charset=utf-8"); simply sends a utf-8 header to ensure the output is encoded with UTF-8. The second line, $str = '我爱编程', is a Chinese characters string. It is 4 Chinese characters. In Chinese, 1 character is 1 word as well, so the string is also 4 words. Translating it into English is 'i love programming', which is 3 words, 18 characters including space.
Now, I want to return part of the Chinese string, starting from position 2, and length is 2. The correct result should be 编程.
The third line: echo substr($str, 2,2), '<br>'. Here we use normal substr. As we can expect, the output would be wrong. On my screen, the output is ��
Next, the last line: echo mb_substr($str, 2, 2), '<br>'. Now we use PHP's mb_substr and expect it could work properly. Does it work? Unfortunately, it doesn't! The output is still ��. Let's check PHP manual about mb_substr: "string mb_substr ( string $str , int $start [, int $length [, string $encoding ]] ). The encoding parameter is the character encoding. If it is omitted, the internal character encoding value will be used." So, based on the manual, we change our code to:
<?php
header("Content-Type:text/html;charset=utf-8");
$str = '我爱编程';
echo substr($str, 2, 2), '<br>';
echo mb_substr($str, 2, 2, 'UTF-8'), '<br>';
?>
This time, it works! The output is 编程, as what we expect. So in this case, we can't simply use mb_substr and expect it can work properly. We still have to specify the encoding method. If we don't specify the character encoding, PHP will use the internal character encoding value. Then we have another question: what is the internal character encoding value? To answer this question, we must have a look at our php configuration file. Let's open our php.ini. We can find a [mbstring] section.
[mbstring]
...
;mbstring.internal_encoding = EUC-JP
...
Ok, that is quite clear now. Let's uncomment this line and change the value to UTF-8, and then restart Apache server(Don't forget this).
[mbstring]
...
mbstring.internal_encoding = UTF-8
...
Now, let's try this code again:
<?php
header("Content-Type:text/html;charset=utf-8");
$str = '我爱编程';
echo substr($str, 2, 2), '<br>';
echo mb_substr($str, 2, 2), '<br>';
?>
Now it works, mb_substr is using internal encoding value, and that value has been set to UTF-8.
One thing i don't like about PHP(and javascript) is, for a same task, it always provides a few different ways to do it. I used to call it too much flexibility (http://hengrui-li.blogspot.com/2011/04/too-much-language-flexibility-good-or.html). Recently i learned the core philosophy of Python: "There should be one - and preferably only one - obvious way to do it" and i found why i don't think PHP is a great programming language(purly from programming language perspective, not from the point that how it boosts web and makes web programming so easy).
So, let's suppose we only want to use one function for the task "to return part of a string". Obviously, mb_substr is our choice, and replace all substr in old system with mb_substr is not hard. But what if we don't want to change our code? Or what if we simply think typing mb_substr is less efficient than typing substr?
Actually, mbstring supports overloading the existing string manipulation functions. If we enable overloading, when we call substr(), PHP will actually call mb_substr() automatically. Let's see how to enable overloading in php.ini:
[mbstring]
...
; overload(replace) single byte functions by mbstring functions.
; mail(), ereg(), etc are overloaded by mb_send_mail(), mb_ereg(),
; etc. Possible values are 0,1,2,4 or combination of them.
; For example, 7 for overload everything.
; 0: No overload
; 1: Overload mail() function
; 2: Overload str*() functions
; 4: Overload ereg*() functions
;mbstring.func_overload = 0
...
So we simply uncomment the last line, and set to value to 7: mbstring.func_overload = 7, restart apache, and try the code:
<?php
header("Content-Type:text/html;charset=utf-8");
$str = '我爱编程';
echo substr($str, 2, 2), '<br>';
echo mb_substr($str, 2, 2), '<br>';
?>
We can find both functions work fine! But doing this overloading can cause issues. If we are using normal string manipulation functions to handle real binary data(it means real binary data, NOT the text string treated as binary), enable overloading could break the binary handling code. Although i hardly see PHP code need to handle real binary data, it is safest that we simply use mb_ string functions. Just remember this on PHP manual also: "It is not recommended to use the function overloading option in the per-directory context, because it's not confirmed yet to be stable enough in a production environment and may lead to undefined behaviour."
Tuesday, August 16, 2011
Selenium user interface test with PHPUnit
I used to simply use Selnium IDE, a firefox plugin, for a website's interface test. However, if we want to integrate the user interface tests into our continuous integration system, we can use Selenium RC server to do automated user interface tests in our continuous integration system. Selenium runs all the tests directly in a browser, just as a real user is browsing the website. PHPUnit provides the functions we need to talk to Selenium RC server and we can write user interface test cases just like usual unit test cases in PHPUnit.
First we must download Selenium Server from here: http://seleniumhq.org/download/, my current version is the latest 2.3.0. It is a single .jar file: selenium-server-standalone-2.3.0.jar. To start the server, we must run this command:
java -jar /path/to/selenium-server-standalone-2.3.0.jar
That is it. We have setup our testing server.
Now, write our first test case, we just want to browse http://localhost/ and check if it works! :D
TestLocalhost.php
<?php
class TestLocalhost extends PHPUnit_Extensions_SeleniumTestCase
{
protected function setUp()
{
$this->setBrowser("*firefox");
$this->setBrowserUrl("http://localhost/");
}
public function testLocalhost()
{
$this->open("/");
$this->verifyTextPresent("It works!");
}
}
?>
And then we can just run the test with the command:
phpunit TestLocalhost.php
Let's check through this test case.
In setUp method, we set up the browser we want to use: $this->setBrowser("*firefox"). This tells selenium to use firefox for testing. We can setup other browsers as long as they are installed on our system, some of them are:
*firefox
*chrome
*iexplore
*safari
*opera
What if we want to run the tests on a series of browsers? We can declare a public static $browsers array in the test class:
class TestLocalhost extends PHPUnit_Extensions_SeleniumTestCase
{
public static $browsers = array(
array(
'name' => 'Firefox on Linux',
'browser' => '*firefox',
'host' => 'localhost',
'port' => 4444,
'timeout' => 30000,
),
array(
'name' => 'Chrome on Linux',
'browser' => '*chrome',
'host' => 'localhost',
'port' => 4444,
'timeout' => 30000,
),
);
protected function setUp()
{
$this->setBrowserUrl("http://localhost/");
}
public function testLocalhost()
{
$this->open("/");
$this->verifyTextPresent("It works!");
}
}
?>
Now, Selenium will run the tests through all browsers declared in the static $browsers array.
$this->setBrowserUrl("http://localhost/"); set up the base Url of our web application.
Now we have one test case testLocalhost(). $this->open("/") tells we open the root of our web site first: http://lcoalhost/. $this->verifyTextPresent("It works!") will verify the text 'It works' should be presented after we go to http://lcoalhost/,
Now let's look at another example:
<?php
class TestWebApp extends PHPUnit_Extensions_SeleniumTestCase
{
protected function setUp()
{
$this->setBrowser("*chrome");
$this->setBrowserUrl("http://webapp/");
}
public function testLogin()
{
$this->open("/");
$this->type("id=txtUserId", "username");
$this->type("id=txtPassword", "wrongpassword");
$this->click("id=frmLoginButton");
$this->waitForPageToLoad("30000");
$this->verifyTextPresent("Username or Password incorrect");
$this->type("id=txtPassword", "correctpassword");
$this->click("id=frmLoginButton");
$this->waitForPageToLoad("30000");
$this->assertEquals("WebApp 3.3.0", $this->getTitle());
}
}
?>
This simple test case test login of a web application. $this->type("id=txtUserId", "username") will type the text 'username' into the text field with id=txtUserId, and then type the text 'wrong password' into the password text field. $this->click("id=frmLoginButton") will do a click action on the login button. $this->waitForPageToLoad("30000"), well, quite self explained. Since we enter a wrong password, we expect the text "Username or Password incorrect" is displayed on the web page, so we do $this->verifyTextPresent("Username or Password incorrect");
Writing all these test cases for a whole web site is quite time consuming. We can use Selenium IDE to record all our tests and export them in PHPUnit format, which makes our life much easier.
Monday, August 15, 2011
MySql stored procedure commands
1. To show all stored procedures of a database:
SHOW PROCEDURE STATUS where DB = 'databasename';
2. Dump(export) all stored procedures of a database
mysqldump -uuser -ppassword --routines databasename > outputfile.sql
Sunday, August 14, 2011
an interesting SQL query question
A very interesting SQL statement question. We have a 'payment' table:
payment table:
year salary
2000 1000
2001 2000
2002 3000
2003 4000
The question is: Write a query on payment table so we get the result below:
Query Result:
year salary
2000 1000
2001 3000
2002 6000
2003 10000
We can find that the result's salary is the sum of the salary of the year and the salary of last year. For example:
In Query Result, 2001's salary is 3000 = Table payment's year 2001's salary 2000 + Table payment's year 2000's salary 1000
Answer 1:
select year, (select sum(b.salary) from payment b where b.year <= a.year) as salary from payment a;
To expalin this one, better imagin we have two tables payment a, and payment b
payment a:
Year payment
2000 1000
2001 2000
2002 3000
2003 4000
payment b:
Year payment
2000 1000
2001 2000
2002 3000
2003 4000
First, MySql tries to retrieve all the year data from table a. For the second column, MySql will select the rows from table b that b.year <= a.year, and then get the SUM of the payment of these rows. The process is like:
For a.2000: b.2000 <= a.2000; SUM(b.1000)
For a.2001: b.2000 <= a.2001; b.2001 <= a.2001; SUM(b.1000,b.2000)
For a.2002: b.2000, b.2001, b.2002 <= a.2002; SUM(b.1000,b.2000,b.3000)
...
Finally, we can get the result:
year salary
2000 1000
2001 3000
2002 6000
2003 10000
Answer 2:
Answer 1 using sub query is straight forward and easier for people to understand. But when it comes to SQL, sub query is mostly not the best option. We can always write the easier sub query first, and then after we fully understand how the data should be retrieved, we can change it to use JOIN. That is how to use JOIN to do it:
select a.year, sum(b.salary) salary from payment b join payment a on b.year <= a.year group by a.year;
Answer 3 (Here is the real interesting answer):
If we still remember the math we learned at primary school(Well, at least that is being taught in primary school in China), we can find that the salary in payment table is actually an arithmetic sequence (http://en.wikipedia.org/wiki/Arithmetic_series)! For arithmetic sequence, we have this formula: Sn = (A1 + An) * n / 2. So, we can use this query:
select year, (1000+salary)*salary/2000 from payment;
It works perfectly for this specific question.
Thursday, August 11, 2011
how to reverse a string in php
This is just a simple question for fun. Well, it could be an interview question. Let's assume we are in an interview and being asked this question. The answers i can come up with are listed below:
$str = 'this is testing'; How to reverse the string to 'gnitset si siht'?
1. strrev, as long as we know this function exists.
$str = 'this is testing';
echo strrev($str);
2. What if we don't know this strrev function? Then this is not bad as well (i think):
$str = 'this is testing';
$strArray = str_split($str);
$reverseArray = array_reverse($strArray);
$reverseStr = implode($reverseArray);
3. If we cannot remember any of these functions, we have to solve this problem by figuring out our own algorithm.
$str = 'this is testing';
$length = strlen($str);
$reverseStr = '';
for($i=$length-1; $i>=0; $i--) {
$reverseStr .= $str[$i];
}
Oh yes, we can access a string's elements via array operator. It is not surprising if you know how C handles string.
4. What if we can't even remember strlen() either? Seriously? Ok, still not a problem:
$str = 'this is testing';
$reverseStr = '';
$i = 0;
while(isset($str[$i])) {
$reverseStr = $str[$i] . $reverseStr;
$i++;
}
As long as we know how PHP handles string internally(actually, how C handles string), we can solve the problem.
Frankly speaking, if i were asked this question in an interview, i could probably only come up with answer 3 and answer 4, cause I can't remember all those PHP functions either.
5. What if we can't even remember isset() function? You must be kidding.
5. What if we can't even remember isset() function? You must be kidding.
6. What if we can't figure out our own algorithm? humm... that is not good. If an interviewer ask you this question, he probably expect you can come up with your own solutions without using any special functions to see your capability to analyse and solve problems(Well, but i do see interview just trying to test candidates' memory). He probably even tells you that you cannot use any of those special functions. Anyway if we really cannot figure out any implementation, I probably would just write down my ultimate answer/solution: Google
:D :D :D
Wednesday, August 10, 2011
php isset, is_null and ===null
This is post is simply for fun.
what is the difference between PHP's isset and is_null? Let's see what PHP manual states:
isset: Determine if a variable is set and is not NULL.
is_null: Finds whether the given variable is NULL.
It seems they work for exactly same purpose simply in an opposite way:
$variable = null;
isset($variable) !== is_null($variable)
or
!(isset($variable)) === is_null($variable)
Is that simple? Not really. One difference between isset and is_null is: as its name suggests, when using isset to check an undefined variable, PHP won't raise a NOTICE. But if we use is_null to check an undefined variable, a NOTICE will be raised.
//a NOTICE will be given by PHP
is_null($undefinedVar);
//no NOTICE
isset($undefinedVar);
In PHP manual, there is a note for isset:
Note: Because this (isset) is a language construct and not a function, it cannot be called using variable functions
This tells us isset is a language construct like echo, for, foreach. It is not a function. So here is one difference
//this works
$func = 'is_null';
$func($variable);
//this doesn't work
$func = 'isset';
$func($variable);
Also, is_null is a function, so it can take a function return value as its argument, while isset cannot do that.
//this works
is_null(getVariable());
//this doesn't work
isset(getVariable());
Since isset is language construct while is_null is a function, you can guess there is performance difference between them. Using language construct is more efficient than calling a function. To check the difference, I will use VLD to expose the opcode(about VLD, check this http://hengrui-li.blogspot.com/2011/07/review-php-opcode-with-vld.html)
test1.php
<?php
$name = null;
isset($name);
?>
php -dvld.active=1 test1.php
line # * op fetch ext return operands
---------------------------------------------------------------------------------
2 0 > EXT_STMT
1 ASSIGN !0, null
3 2 EXT_STMT
3 ZEND_ISSET_ISEMPTY_VAR 5 ~1 !0
4 FREE ~1
4 5 > RETURN 1
As we can see, PHP parse isset as one opcode operation: ZEND_ISSET_ISEMPTY_VAR.
test2.php
<?php
$name = null;
is_null($name);
?>
php -dvld.active=1 test2.php
line # * op fetch ext return operands
---------------------------------------------------------------------------------
2 0 > EXT_STMT
1 ASSIGN !0, null
3 2 EXT_STMT
3 EXT_FCALL_BEGIN
4 SEND_VAR !0
5 DO_FCALL 1 'is_null'
6 EXT_FCALL_END
4 7 > RETURN 1
To call a is_null function, PHP does these opcode operations: EXT_FCALL_BEGIN, SEND_VAR, DO_FCALL, EXT_FCALL_END
The result is quite obvious now. isset is more efficient. But sometimes, checking is null simply makes more sense than checking isset, especially when we are checking a function return value:
$name = $user->getName();
isset($name);
is_null($name);
For this example, we can still use isset($user) to check if $user is null or not, but, by the meaning of the name 'isset', here $user seems being set obviously. What we really want to check is if $user is null or not. At this situation(when we simply want to check if a variable is null, not worrying about if it is set), using $name === null is better than using is_null($name)
$name === null returns exactly same result as is_null($name), however, $name === null is more efficient. It is almost as efficient as isset. Let's check the opcode:
test3.php
<?php
$name = null;
$name === null;
?>
php -dvld.active=1 test3.php
line # * op fetch ext return operands
---------------------------------------------------------------------------------
2 0 > EXT_STMT
1 ASSIGN !0, null
3 2 EXT_STMT
3 IS_IDENTICAL ~1 !0, null
4 FREE ~1
4 5 > RETURN 1
From micro performance optimization perspective: isset better than ===null better than is_null.
I think the correct usage of them is: if we want to check if a variable is "set"(or existing), use isset; if we want to check if an existing variable's value is null, use === null.
The performance difference among them belongs to micro-optimization, so i have to run 5 million times of comparison to see the difference:
The performance difference among them belongs to micro-optimization, so i have to run 5 million times of comparison to see the difference:
My testing code is quite simple:
<?php
$counter = 5000000;
$name = 'henry';
$start = microtime(true);
for($i=0; $i<$counter; $i++) {
isset($name);
}
$end = microtime(true);
echo $end - $start, "\n";
?>
?>
I simply replace isset(name) with is_null($name) and $name === null and run the script respectively. And the result is:
isset: 1.62s
===null: 1.80s
is_null: 8.47s
It is a little out of my expect that is_null could be that slower, anyway, just a simple test for fun.
Subscribe to:
Posts (Atom)